A method for locating a fire in a large-span space warehouse shelf

By combining the fire detection and shelf positioning methods of YoloV3 and HED algorithms, the delay and inaccurate problems of fire detection and positioning in large-span space buildings are solved, and the precise positioning of the fire point of the shelf is achieved, and the control effect of the fire extinguishing device is improved.

CN117152663BActive Publication Date: 2025-07-22TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311186986.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-14
Publication Date
2025-07-22
Estimated Expiration
2043-09-14

AI Technical Summary

Technical Problem

There are problems of delay and inaccurate fire detection and positioning in large-span buildings, especially the poor three-dimensional positioning effect of the shelf fire point, which leads to poor control effect of the fire extinguishing device.

Method used

The fire detection module based on YoloV3 object detection algorithm and the shelf positioning module based on the HED edge detection algorithm are adopted, and the fire detection is combined with Focal and IOU loss functions, and the shelf row division is divided by cross entropy and Dice loss functions to achieve precise positioning.

Benefits of technology

It realizes accurate positioning of shelves in large-span warehouses, improves the accuracy of fire detection and the identification accuracy of shelf ranks, and supports accurate fire extinguishing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117152663B_ABST
    Figure CN117152663B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for locating a fire in a large-span space warehouse shelf, comprising the following steps: Video input: Reading the monitoring video frames of the warehouse shelf and performing image preprocessing operations including filtering, noise reduction, and correction; Fire detection: Detecting whether a fire has occurred by passing the preprocessed video frames through a fire detection module, and if so, outputting pixel-level fire coordinates; During the training process of the fire detection network, a form of fusing the Focal loss function and the IOU loss function is used to train the network; Shelf row and column division: Dividing the shelves in the normal state through a shelf positioning module, and the shelf positioning module is constructed based on the HED edge detection algorithm; Mapping the pixel-level fire coordinates obtained in step 2 into the shelf row and column contour map obtained in step 3, and outputting the row and column information of the shelf where the fire is located.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and deep learning, and particularly to a method for locating a fire on a large-span space warehouse shelf. Background Art

[0002] With the rapid development of the social economy, there are more and more large-space buildings such as enterprise production, large wholesale markets, storage buildings, and the logistics industry. These buildings have the characteristics of large personnel mobility, a large amount of combustible substances stored, and intricate internal channels. Once a fire occurs, it is extremely easy to cause heavy property losses and major accidents.

[0003] Limited by the internal structure of the large-span space building, the installed temperature sensors and sprinkler systems are far from the fire ignition point, and the sprinkler device will only be activated when the temperature reaches a relatively high threshold, which will undoubtedly lead to untimely fire detection. In addition, due to the large scope of the large-span space, installing so many temperature sensors will consume a huge amount of manpower, material resources, and financial resources.

[0004] The goods stacked on the warehouse shelves are in a dynamic change, and the changing stacking of goods will affect the effect of three-dimensional positioning of the fire on the shelves. Therefore, the fire extinguishing effect of controlling the fire extinguishing device through the pitch angle in a large-span space is poor, and the control effect on a certain range of the fire scene is not good. Summary of the Invention

[0005] The present invention provides a method for locating a fire on a large-span space warehouse shelf with high precision, and the technical solution is as follows:

[0006] A method for locating a fire on a large-span space warehouse shelf includes the following steps:

[0007] Step 1, video input: Read the monitoring video frames of the warehouse shelf and perform image preprocessing operations including filtering, noise reduction, and calibration;

[0008] Step 2, fire detection: Detect whether there is a fire in the video frames after the preprocessing operation through a fire detection module. If there is a fire, output the pixel-level fire coordinates; during the training process of the fire detection network, the network is trained in a form of fusing the Focal loss function and the IOU loss function;

[0009] Step 3, shelf row and column division: Divide the shelves in the normal state into rows and columns through a shelf positioning module. The shelf positioning module is constructed based on the HED edge detection algorithm, and the method is as follows:

[0010] The edge detection network constructed by the HED edge detection algorithm has a backbone network divided into several stages. Each stage consists of 3 convolutional-batch normalization-ReLU groups, and max-pooling is used to connect different stages to form a multi-scale network structure. After each stage in the backbone network, a holistic edge attention SE module is introduced. The SE module is composed of a Laplacian convolutional layer with fixed weights, an average pooling layer, and two fully connected layers connected in series in sequence. After the SE module in each stage of the backbone network, a bypass network is also connected in parallel. The bypass network uses a 1×1 convolutional layer to compress the features extracted by the backbone network. The bypass feature maps of each stage obtained are concatenated by channels to obtain a roughly extracted edge map. The roughly extracted edge map is refined by rows and columns through the non-maximum suppression algorithm to obtain the final shelf row and column contour map.

[0011] During the training process of the edge detection network, a deep supervision mechanism is adopted, that is, the loss is calculated for all the bypass output maps of each stage and the final edge map. The loss function is a weighted fusion of the cross-entropy loss function and the Dice loss function.

[0012] Step 4, coordinate mapping: Map the pixel-level fire coordinates obtained in Step 2 into the shelf row and column contour map obtained in Step 3, and output the row and column information of the shelf where the fire is located.

[0013] Among them, the fire detection network of the fire detection module is constructed based on the YoloV3 object detection algorithm. The backbone network is used to extract image features. After the backbone network, multiple detection heads are introduced, and each detection head is responsible for object detection at different scales. The recognized bounding boxes are filtered by using the non-maximum suppression method, and only the bounding box with the highest confidence is retained to obtain the final fire detection result, and at the same time, the position coordinates of the center point of the bounding box are output.

[0014] The method of applying the edge detection algorithm to the fire location in the special scenario of large-span space warehouse shelves in the present invention has high practical value and economic value, and provides a new idea for intelligent fire extinguishing. Description of the Drawings

[0015] Figure 1 is the flowchart of the method of the present invention;

[0016] Figure 2 is the network structure diagram of YoloV3;

[0017] Figure 3 is the network structure diagram of HED+SE;

[0018] Figure 4 is the high-definition picture of the warehouse shelf in the normal state;

[0019] Figure 5It is the fire detection result diagram of the method of the present invention;

[0020] Figure 6 It is the fire location result diagram of the method of the present invention. Specific embodiments

[0021] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0022] This embodiment provides a method for locating a fire in a large-span space warehouse shelf, and the specific steps are as Figure 1 shown, including the following steps:

[0023] (1) Read video frames from a local disk, a hard disk video recorder or a network camera to complete data acquisition, and then perform preprocessing operations such as filtering and correction on the collected pictures. Assuming the input picture is I and the output picture is O, the process of image preprocessing can be described by the following formula:

[0024] O = median[rect(I)]

[0025] In the above formula, median represents the use of the median filtering algorithm. The specific process is to create a template of size 3×3, and then sort the pixels in the template according to the pixel value size to generate a two-dimensional data sequence that increases (or decreases) monotonically. Then, replace all pixel points in this area with the median pixel, and finally slide the template until all pixel points are processed. rect represents the image correction algorithm, which restores the distorted image to a normal state. After preprocessing, both image noise and distortion are eliminated, which is beneficial to the processing of subsequent steps.

[0026] (2) Detect whether a fire has occurred in the preprocessed video frames through the fire detection module. If so, output the fire pixel-level coordinates. Among them, the fire detection module is constructed based on the YoloV3 object detection algorithm, and the detection speed is 3fps, meeting the real-time requirement. The network structure diagram is as Figure 2 shown. During the detection process, first, the image needs to be converted into a 4D tensor tensor = [N, C, H, W]. N represents the number of pictures processed at one time, which is set to 1 here; C represents the number of picture channels. Here, the read image is in RGB representation, so C = 3; H and W respectively represent the height and width of the picture. After bilinear interpolation, the picture scale is uniformly scaled to 1080×720. Therefore, the obtained 4D tensor tensor = [1, 3, 400, 1080].

[0027] The YoloV3 object detection algorithm is very efficient. It can predict the classes and locations of multiple objects simultaneously in a single forward pass. Specifically, YoloV3 uses Darknet-53 as the backbone network. This network mainly consists of 1 DBL module and 5 residual connection units Res_n. The DBL module is composed of 1 convolutional layer, 1 batch normalization layer, and 1 leaky-relu activation function. The residual connection unit is composed of 2 DBL modules combined with a shortcut connection. n represents the number of residual connection units included. Therefore, the structure of Darknet-53 is DBL-Res1-Res2-Res8-Res8-Res4. This backbone network is used to extract image features. After Darknet-53, 3 detection heads are introduced. Each detection head is responsible for object detection at different scales. Each detection head is sequentially connected by 5 DBL modules and 1 ordinary convolutional layer, aiming to further extract image features. By introducing detection heads of different scales, the YoloV3 object detection algorithm can better adapt to multi-scales and improve detection performance. Finally, the identified bounding boxes are filtered using the non-maximum suppression method, and only the bounding boxes with the highest confidence are retained to obtain the final fire detection result, and the position coordinates of the center points of the bounding boxes are output simultaneously.

[0028] The data pictures of the fire detection network are generated by intercepting video frames of the burning shelf, including a total of 2500 pictures, of which 2000 are used for training and 500 are used for testing. All pictures are labeled for the flame using a labeling tool, with a resolution of 400×1080 and arranged in a unified naming manner.

[0029] During the training process, the network is trained in the form of fusing the Focal loss function and the IOU loss function. In the fire detection task, the fire area usually occupies a small part of the image, while the non-fire area occupies most of it. This class imbalance may cause the model to tend to ignore the fire area and focus on the background. By introducing the Focal loss, the model can pay more attention to difficult samples, that is, those fire areas that are difficult to classify. The IOU loss function is very useful in the regression task of the fire location. It emphasizes the accurate positioning of the fire and helps to ensure the accurate location of the detection result. After fusing the IOU loss function, the network can better predict the position of the target bounding box, thereby improving the positioning accuracy of the fire detection module. Therefore, fusing the Focal and IOU loss functions helps to further improve the detection performance of the fire detection module. The specific formula is as follows:

[0030]

[0031] In the above formula, L cls represents the Focal loss function, and L regRepresents the IOU loss function, N pos represents the number of positive samples, and λ = 1 is the balance weight of L reg is an indicator function. If C * > 0, it is 1; otherwise, it is 0.

[0032] The overall training process uses the PyTorch deep learning framework. The optimization method is Adam. The minimum data batch is set to 2, the number of training epochs is 40, the initial learning rate is 0.0001, and the learning rate is reduced by 10 times every 5 epochs.

[0033] (3) The shelf positioning module is constructed using the HED-based edge detection algorithm. The network structure is as Figure 3 shown. During the row and column division process, it is also necessary to convert the image data into a 4D tensor, and the tensor size is tensor = [1, 3, 300, 400].

[0034] The HED edge detection algorithm has the advantages of high efficiency and good accuracy. The backbone network is divided into 5 stages, and each stage consists of 3 convolutional-batch normalization-ReLU groups. Different stages are connected by max pooling to form a multi-scale network structure. To suppress the background information of the shelf warehouse and enhance the ability of the edge detection algorithm to extract the shelves' rows and columns, a holistic edge attention SE module is introduced after each convolutional layer in the backbone network. This module is composed of a Laplacian convolutional layer with fixed weights, an average pooling layer, and two fully connected layers connected in series in sequence. This improvement makes the HED edge detection algorithm more adaptable to the special scenario of large-span space warehouse shelves. A bypass network is also connected in parallel after the backbone network. This sub-network uses a 1×1 convolutional layer to compress the features extracted by the backbone network, and then the 5 bypass feature maps obtained are concatenated by channels to obtain a roughly extracted edge map. Finally, this edge image is refined by the non-maximum suppression algorithm to obtain the final shelf row and column contour map.

[0035] The data images of the edge detection network are collected by taking pictures of the non-fire shelves from four angles: front, back, left, and right. There are a total of 200 images. Then, the dataset is expanded through data augmentation means such as cropping, rotation, and distortion, and finally expanded to 1000 images, among which 800 are used for training and 200 are used for testing. All images are labeled with the shelf row and column contours using a labeling tool, with a resolution of 300×400, and arranged in a unified naming manner.

[0036] ​The training process adopts a deep supervision mechanism, that is, the losses are calculated for all 5 bypass output images and the final edge image. The loss function is a weighted fusion of the cross-entropy loss function and the Dice loss function. The fusion of these two loss functions can improve the network's ability to locate the shelf contour. The cross-entropy loss function is usually used for image pixel-level classification, which helps the network correctly classify the pixels of the shelf, while the Dice loss function focuses on the accuracy of the shelf boundary, improving the accuracy of boundary detection. By using these two loss functions through weighted fusion, the network can not only accurately classify the shelf row and column pixels but also better capture the row and column boundary information of the shelf. Therefore, it is easier to extract a more refined contour of the shelf. The specific formula is as follows:

[0037]

[0038]

[0039] loss fusion = loss cross-entropy + loss Dice

[0040] In the above formula, loss cross-entropy represents the cross-entropy loss function, loss Dice represents the Dice loss function, loss fusion represents the fused loss function, y i and l i represent the i-th pixel point in the output image Y and the label image L respectively, and n represents the number of all pixel points. The network is built using the PyTorch deep learning framework. The optimization method is Adam, the minimum data batch is set to 8, the number of training epochs is 30, the initial learning rate is 5e-4, and the learning rate is reduced by 10 times every 5 epochs.

[0041] (4) Map the pixel-level coordinates of the fire ignition point output in step (2) to the shelf row and column contour map in step (3) according to the equal-proportion mapping relationship, and output the row and column information at the same time. The positioning result is as Figure 5 shown, and then call the sprinkler device to accurately extinguish the fire at the ignition point.

Claims

1. A method for locating a fire in a large-span space warehouse shelf, comprising the following steps: Step 1, video input: Read the monitoring video frames of the warehouse shelf and perform image preprocessing operations including filtering, noise reduction, and correction; Step 2, fire detection: Detect whether there is a fire in the preprocessed video frames through a fire detection module. If there is a fire, output the pixel-level fire coordinates; during the training process of the fire detection network, a form of fusing the Focal loss function and the IOU loss function is used to train the network; Step 3, shelf row and column division: Divide the shelves in the normal state through a shelf positioning module. The shelf positioning module is constructed based on the HED edge detection algorithm, and the method is as follows: The edge detection network constructed by the HED edge detection algorithm has a backbone network divided into several stages. Each stage consists of 3 convolutional-batch normalization-ReLU groups, and different stages are connected by maximum pooling to form a multi-scale network structure; after each stage in the backbone network, a global edge attention SE module is introduced. The SE module is composed of a Laplacian convolutional layer with a fixed weight, an average pooling layer, and 2 fully connected layers connected in series in sequence; after the SE module of each stage in the backbone network, a bypass network is also connected in parallel. The bypass network uses a 1×1 convolutional layer to compress the features extracted by the backbone network; the bypass feature maps of each stage obtained are concatenated by channels to obtain a roughly extracted edge map; the roughly extracted edge map is refined by rows and columns through the non-maximum suppression algorithm to obtain the final shelf row and column contour map; During the training process of the edge detection network, a deep supervision mechanism is adopted, that is, the loss is calculated for all the bypass output maps of each stage and the final edge map. The loss function is a weighted fusion of the cross-entropy loss function and the Dice loss function; Step 4, coordinate mapping: Map the pixel-level fire coordinates obtained in Step 2 to the shelf row and column contour map obtained in Step 3, and output the row and column information of the shelf where the fire is located.

2. The large-span space warehouse shelf fire location method according to claim 1, characterized in that, Among them, The fire detection network of the fire detection module is constructed based on the YoloV3 object detection algorithm. The backbone network is used to extract image features; after the backbone network, multiple detection heads are introduced. Each detection head is responsible for object detection at different scales; the identified bounding boxes are filtered by using the non-maximum suppression method, and only the bounding box with the highest confidence is retained to obtain the final fire detection result, and at the same time, the position coordinates of the center point of the bounding box are output.

Citation Information

Patent Citations

  • Elevated warehouse intelligent warehouse management system based on unmanned aerial vehicle and management method thereof

    CN110963034A

  • Large-span space warehouse shelf fire positioning control method and system

    CN115601909A