A Fire Detection and Early Warning Method and System Based on Improved YOLOv8

By introducing the SMD-neck structure and SPD-Conv module in the YOLOv8 network, the problem of insufficient target detection accuracy at small scales and high model calculation complexity is solved, and high precision and lightweight fire detection effect is achieved.

CN120047759BActive Publication Date: 2025-07-01XIDIAN UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510536345.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-01
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

The prior art has low resolution and small-scale problems in fire detection. Traditional object detection algorithms are not effective in detecting small-scale flames and smoke, and the computing power of the model is limited on edge devices, reducing real-time performance.

Method used

By introducing MLCA mixed local channel attention and DySample upsampler to improve the Slim-neck structure, the SMD-neck structure was designed to replace the neck structure of the original YOLOv8 network; at the same time, the SPD-Conv module was introduced to replace the convolutional layer in the backbone network, and the feature extraction and calculation complexity of the model were optimized.

Benefits of technology

It significantly improves the detection accuracy of small-scale flame targets, reduces the amount of model parameters and calculation complexity, reduces the demand for edge equipment performance, and achieves high-precision and lightweight fire detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047759B_ABST
    Figure CN120047759B_ABST
Patent Text Reader

Abstract

A fire detection and early warning method and system based on improved YOLOv8 belongs to the field of image recognition technology. The method includes: establishing a fire dataset, annotating it, and dividing it into a training set, a validation set, and a test set; improving the YOLOv8 network; training, evaluating, adjusting, and testing the original YOLOv8 network and the improved YOLOv8 network to determine the optimal model; deploying the optimal model to the edge AI development board of the ground station; obtaining the real-time video stream of the drone through the drone camera and transmitting it to the edge AI development board for fire detection; using the edge AI development board to detect the fire signal and giving an early warning through the fire alarm module. The system includes: a fire dataset establishment and processing module, a YOLOv8 network improvement module, a YOLOv8 network training module, an optimal model deployment module, a fire detection module, and a fire early warning module. The present invention has the technical advantages of high precision and lightweight.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image recognition, and mainly relates to a fire detection and early warning method and system based on improved YOLOv8. Background Art

[0002] Traditional fire monitoring schemes mainly rely on various sensors, image processing technologies, and manual inspections to achieve fire detection and early warning.

[0003] Sensor detection technology: In traditional fire monitoring schemes, sensor technology is the foundation, mainly including temperature sensors, smoke sensors, and flame sensors, etc. These sensors can monitor the temperature, smoke concentration, and the presence of flames in the environment, thereby achieving a preliminary judgment of the fire situation. However, limited by the sensor layout, only fixed areas can be monitored, and large-scale coverage detection cannot be achieved.

[0004] Manual inspection: Traditional fire monitoring also relies on methods such as manual mountain patrols, setting up observation towers, or remote aerial photography. These methods require a large amount of labor costs, and it is difficult to detect early fires, which brings difficulties to fire extinguishing.

[0005] Image processing technology: In some traditional schemes, digital image processing technology is also used to identify fire signs, such as smoke and flames. This method depends on feature selection and calculations in the preprocessing stage, but in practical applications, there may be problems such as insufficient accuracy, weak adaptability, and low detection efficiency.

[0006] Satellite remote sensing technology: Satellite remote sensing technology can cover a vast area and has a certain detection ability for fires in open areas such as forests and farmlands. However, the response speed of this method is restricted by the satellite overpass cycle, and its resolution may not be sufficient to capture small-scale fire sources.

[0007] In recent years, unmanned aerial vehicle (UAV) technology has been widely applied in many fields, and fire inspection and early warning is an important direction. Its main advantage is that UAVs can quickly cover a vast area and provide high-resolution video streams for real-time fire detection. In addition, using object detection technology based on deep learning, fire signs can be accurately detected in the video stream.

[0008] However, applying UAV video streams to fire detection is not easy. Traditional object detection algorithms such as Faster R-CNN, SSD, etc., although performing excellently in general object detection tasks, may not be able to accurately and real-time detect flame and smoke targets in special scenarios like fires, especially from the remote sensing perspective with low resolution and high light flux density differences.

[0009] The YOLO (You Only Look Once) algorithm is a popular real-time object detection system. Its core idea is to solve the object detection task as a regression problem. Different from traditional object detection methods, the YOLO algorithm can predict the positions and classes of all objects in an image with only one forward pass. The YOLOv8 algorithm is the latest generation of object detection algorithm in the YOLO series, which provides cutting-edge performance in terms of accuracy and speed.

[0010] The patent application document with the publication number CN116206223A proposes a fire detection method and system based on drone edge computing. The invention obtains the detection image of the target area by taking pictures with a drone, and uses the Yolov3 algorithm for deep learning edge detection processing, which can quickly and accurately detect the position information of the fire ignition point, and has intelligence and high efficiency.

[0011] The patent application document with the publication number CN114037910A discloses a drone forest fire detection system, including: a drone, an image processing unit, a control unit and a fire alarm. Among them, the drone is responsible for collecting image data of the forest area and transmitting it to the image analysis module; the image analysis module uses the YOLOv7 algorithm to analyze the collected images to identify possible fire signs and transmits the analysis results to the control center; the control center judges whether a fire has occurred based on the received analysis results. Once a fire is confirmed, an activation signal is sent to the fire alarm device; after receiving the signal, the fire alarm device immediately starts the alarm program and issues a fire alarm. The system aims to improve the detection accuracy of early fires.

[0012] The patent application document with the publication number CN116229296A discloses a fire detection system using deep learning technology, including a drone and a fire detection algorithm based on the YOLOX framework. The system captures environmental images through the drone, analyzes the images using the algorithm to identify fires, and automatically issues an alarm when a fire is detected, thereby improving the efficiency and response speed of fire early warning.

[0013] The patent application document with the publication number CN117409191A proposes a fire warning method combining a drone and an improved YOLOv8 algorithm. First, collect and label fire data, then use this data to train the original YOLOv8 algorithm, and optimize it using the Omni-Directional Dynamic Convolution (ODConv) and channel coordinated attention mechanism. The optimized model is deployed on a NVIDIA development board for real-time analysis of the drone video stream to detect fire signs. Once a fire is detected, the system will automatically issue a warning. This method improves the accuracy and stability of fire detection, realizes the efficient cooperation between the drone and the ground system, and provides a faster and more accurate fire warning solution.

[0014] The patent application document with publication number CN117333753A proposes a fire detection method based on PD - YOLO. The main steps include: First, preprocess the fire dataset and split it into a training set and a test set; Second, construct and improve the YOLOv8 network model, introducing PConvs and DYDPConv modules to replace some C2f modules; Finally, use the training set to train the improved model and apply it to fire detection. This method enhances the model's learning ability for flame and smoke features, improves the effects of feature extraction and multi - scale feature fusion, and thus more effectively identifies key features related to fires.

[0015] In summary, the existing technologies have the following problems:

[0016] 1. Low - resolution and small - scale problems: During high - altitude inspections, the early fire features obtained by drones or in images will become very small. In this case, traditional object - detection algorithms may not perform well in detecting small - scale flames and smoke. Existing methods such as CN116206223A, CN114037910A, and CN116229296A, although all using deep - learning technologies, mainly rely on existing YOLO - series algorithms (such as YOLOv3, YOLOX). These algorithms may not be optimized for the specific problems in the fire scene (such as small - scale flames, etc.). Although the patent application document with publication number CN117333753A enhances the model's learning ability for flame and smoke features and improves the effects of feature extraction and multi - scale feature fusion, it does not mention the detection effect of the model on small - target flames.

[0017] 2. Lightweight problem: In the fire - detection scenario, a large number of cases involve deploying the model to edge devices for detection. However, the computing power, memory, and storage space of edge devices are usually limited. For traditional YOLO - series algorithms, it is difficult to achieve fast inference, which greatly reduces the real - time performance of fire monitoring. The patent application document with publication number CN117409191A, although having certain improvements in small - scale flame detection, does not perform lightweight improvement on the model, and has too high requirements for the performance of the deployed devices. Summary of the Invention

[0018] To overcome the deficiencies of the above-mentioned existing technologies, the purpose of the present invention is to provide a fire detection and early warning method and system based on improved YOLOv8. By introducing the MLCA hybrid local channel attention and the DySample upsampler to improve the Slim-neck structure, the designed SMD-neck structure is used to replace the neck structure of the original YOLOv8 network. At the same time, the SPD-Conv module is introduced to replace the second, third, fourth, and fifth convolutional layers in the backbone network of the original YOLOv8 network to improve the original YOLOv8 network. Then, the improved YOLOv8 network is trained and tested, etc. The selected optimal model is mounted on a drone, and the drone aerial photography is used to monitor the initial fire in real time. It can improve the detection performance of small target flames while realizing model lightweight and reducing the performance requirements for deployment devices. The present invention has the technical advantages of high precision and lightweight.

[0019] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0020] A fire detection and early warning method based on improved YOLOv8, comprising the following steps:

[0021] S100: Establish a fire dataset, annotate the data in the fire dataset, and divide it into a training set, a validation set, and a test set;

[0022] S200: Improve the original YOLOv8 network to obtain an improved YOLOv8 network;

[0023] S300: Input the training set divided in step S100 into the original YOLOv8 network and the improved YOLOv8 network in step S200 for training respectively; use the validation set divided in step S100 to evaluate and adjust the model performance obtained during the two training processes to optimize the model parameters; input the test set divided in step S100 into the models obtained by training the original YOLOv8 network and the improved YOLOv8 network after being evaluated and adjusted by the validation set for testing to determine the optimal model;

[0024] S400: Deploy the optimal model obtained in step S300 to the edge AI development board of the ground station to obtain the edge AI development board of the ground station with the optimal deployed model;

[0025] S500: Obtain the real-time video stream of the drone through the drone camera and transmit it to the edge AI development board of the ground station with the optimal deployed model in step S400 for fire detection;

[0026] S600: When the edge AI development board of the ground station with the optimal deployed model in step S500 detects a fire signal, give an early warning through the fire alarm module.

[0027] The specific method of step S100 includes:

[0028] S101: Collect flame and smoke images in different weather conditions, different lighting conditions, and different regions to construct a fire dataset;

[0029] S102: Screen the data in the fire dataset constructed in step S101 and delete low-quality data;

[0030] S103: Label the screened fire dataset;

[0031] S104: Randomly divide the fire dataset labeled in step 103 into a training set, a validation set, and a test set.

[0032] The specific method of step S200 includes:

[0033] S2001: Introduce the SPD-Conv module to replace the second, third, fourth, and fifth convolutional layers in the backbone of the original YOLOv8 network; the original YOLOv8 network includes a backbone, a neck structure, and a detection head;

[0034] S2002: Construct an SMD-neck structure to replace the neck structure of the original YOLOv8 network; specifically including:

[0035] Introduce the Slim-neck structure and introduce the MLCA mixed local channel attention in the Slim-neck structure, that is, add the MLCA mixed local channel attention after the second, third, and fourth VoV-GSCSP modules in the Slim-neck structure to obtain the SM-neck structure;

[0036] Introduce the DySample upsampler to replace the Upsample upsampler in the SM-neck structure, that is, replace the first and second Upsample upsamplers in the SM-neck structure with the DySample upsampler to obtain the SMD-neck structure;

[0037] Replace the neck structure of the original YOLOv8 network with the obtained SMD-neck structure.

[0038] In step S2001, introduce the SPD-Conv module to replace the second, third, fourth, and fifth convolutional layers in the backbone of the original YOLOv8 network; the working process of the backbone of the improved YOLOv8 network includes:

[0039] The input feature map passes through the first convolutional layer (Conv), the second SPD-Conv module, the first C2f module, the third SPD-Conv module, and the second C2f module in sequence to obtain the third-layer feature map (P3). The third-layer feature map (P3) is input into the neck structure (Neck). The third-layer feature map (P3) then passes through the fourth SPD-Conv module and the third C2f module in sequence to obtain the fourth-layer feature map (P4). The obtained fourth-layer feature map (P4) is input into the neck structure (Neck). The fourth-layer feature map (P4) then passes through the fifth SPD-Conv module, the fourth C2f module, and the SPPF module in sequence to obtain the fifth-layer feature map (P5). The obtained fifth-layer feature map (P5) is input into the neck structure (Neck).

[0040] In step S2002, an SMD-neck structure is constructed to replace the neck structure (Neck) of the original YOLOv8 network. The workflow of the SMD-neck structure is as follows: The fifth-layer feature map (P5) output by the backbone network (Backbone) of the YOLOv8 network improved in step S2002 passes through the first DySample upsampler and is concatenated with the fourth-layer feature map (P4) output by the backbone network (Backbone) through the first feature fusion layer (Concat) in the channel dimension to obtain a new feature map PA. The feature map PA passes through the first VoV-GSCSP module and the second DySample upsampler in sequence and is concatenated with the third-layer feature map (P3) output by the backbone network (Backbone) through the second feature fusion layer (Concat) in the channel dimension to obtain a new feature map PB. The feature map PB passes through the second VoV-GSCSP module and the first MLCA hybrid local channel attention in sequence to obtain a feature map PC. The feature map PC is input into head-1 of the detection head. After passing through the first GSConv module, the feature map PC is concatenated with the feature map output by the first VoV-GSCSP module through the third feature fusion layer (Concat) in the channel dimension. The output feature map passes through the third VoV-GSCSP module and the second MLCA hybrid local channel attention in sequence to obtain a feature map PD. The feature map PD is input into head-2 of the detection head. After passing through the second GSConv module, the feature map PD is concatenated with the fifth-layer feature map (P5) through the fourth feature fusion layer (Concat) in the channel dimension. The output feature map passes through the fourth VoV-GSCSP module and the third MLCA hybrid local channel attention in sequence to obtain a feature map PE. The feature map PE is input into head-3 of the detection head.

[0041] Replace the neck structure of the original YOLOv8 network with the SMD-neck structure.

[0042] In the step S2001, the working process of the SPD-Conv module includes: the input feature map passes through the spatial-to-depth (SPD) layer, rearranges the spatial blocks of pixels to the depth dimension, then merges different channel groups in the depth dimension, and finally obtains the output feature map through a non-strided convolutional layer;

[0043] In the step S2002, the working process of the GSConv module includes: the input feature map with the number of channels passes through the convolutional layer (Conv) to obtain a feature map with the number of channels, and then passes through the depthwise separable convolutional layer (DWConv) and is concatenated with the feature map with the number of channels output by the convolutional layer (Conv) in the feature fusion layer (Concat). After the concatenated feature map undergoes a shuffle operation, an output feature map with the number of channels is obtained;

[0044] The working process of the VoV-GSCSP module includes: the input feature map sequentially passes through the first convolutional layer (Conv) and the GS bottleneck module, and then is concatenated with the output feature map of the first convolutional layer (Conv) in the feature fusion layer (Concat). The concatenated feature map passes through the second convolutional layer (Conv) to obtain the output feature map;

[0045] The working process of the GS bottleneck module includes: the input feature map passes through the convolutional layer (Conv) to obtain the feature map Pa; the input feature map sequentially passes through two GSConv modules to obtain the feature map Pb; the feature map Pa and the feature map Pb are added to obtain the output feature map;

[0046] The working process of the MLCA hybrid local channel attention includes: the input feature map is first sequentially processed by local average pooling (LAP) and global average pooling (GAP). For the feature map processed by local average pooling (LAP), reshape, 1D convolution (Conv1d), and reshape are sequentially used to obtain the feature map Pc; for the feature map processed by global average pooling (GAP), after 1D convolution (Conv1d), an unpooling (UNAP) operation is performed, and then it is added to the feature map Pc. After the feature map output by the addition operation undergoes an unpooling (UNAP) operation, it is multiplied by the original input feature map to obtain the output feature map;

[0047] The DySample upsampler for the input features Its upsampling process is expressed as:

[0048]

[0049] Wherein, is the sampling set generated by the sampling point generator, is the output; The DySample upsampler generates the sampling set through two types of sampling point generators, static and dynamic;

[0050] The working process of the DySample upsampler includes: The input feature map creates a sampling set through the sampling point generator, and the sampling set and the input feature map are operated through the grid_sample function for resampling to obtain the upsampled feature map;

[0051] The sampling point generator includes a static range factor and a dynamic range factor; Wherein:

[0052] The working process of the static range factor: The input feature map passes through a linear layer and then combines with a fixed range factor (0.25), and then generates an offset ( ) through the pixel shuffle technique, and the offset ( ) is added to the original network position ( ) to obtain the sampling set ( );

[0053] The working process of the dynamic range factor: The input feature map passes through a linear layer and then combines with the dynamic range factor ( ), the obtained output is multiplied by the feature map passing through the linear layer, and then an offset ( ) is generated through the pixel shuffle technique, and the offset ( ) is added to the original network position ( ) to obtain the sampling set ( ), where represents the Sigmoid function for generating the range factor.

[0054] Step S300 includes the following steps:

[0055] S301: Input the training set divided in step S100 into the original YOLOv8 network for training;

[0056] S302: Input the training set divided in step S100 into the improved YOLOv8 network in step S200 for training;

[0057] S303: Compare the training results of steps S301 and S302 to obtain the optimal model. The model parameters for comparison are the number of parameters, the number of floating-point operations (GFLOPs), and the mean average precision (mAP). The specific comparison method is as follows: Introduce the APG coefficient to evaluate the comprehensive performance of the model. The larger this index, the better the performance. The calculation formula for the APG coefficient is:

[0058]

[0059] Among them, 、 、 are weight coefficients used to measure the importance of the three indicators. is the number of parameters of the model obtained by training the original YOLOv8 network. is the number of floating-point operations of the model obtained by training the original YOLOv8 network.

[0060] Step S500 includes the following steps:

[0061] S501: The drone is equipped with a camera, and the video stream captured is transmitted in real time to the 5.8G FPV video transmission receiving module through the 5.8G FPV video transmission module.

[0062] S502: The edge AI development board with the optimal model deployed at the ground station in step S400 receives the real-time video stream of the drone through the 5.8G FPV video transmission receiving module.

[0063] S503: Use the edge AI development board with the optimal model deployed at the ground station in step S400 to detect the video stream received in step S502.

[0064] The said step S600 includes the following steps:

[0065] S601: When the edge AI development board with the optimal model deployed at the ground station detects a fire signal, it conducts serial communication with the STM32C8T6 development board of the fire alarm module through the USART port of the edge AI development board.

[0066] S602: When the STM32C8T6 development board receives the instruction, it communicates with the first LoRa wireless communication module carried on the drone through the second LoRa wireless communication module of the fire alarm module.

[0067] S603: When the first LoRa wireless communication module on the drone receives the instruction, it transmits the longitude and latitude information of the GPS module back to the STM32C8T6 development board of the fire alarm module.

[0068] S604: The STM32C8T6 development board of the fire alarm module uploads the longitude and latitude information and fire images to the fire monitoring cloud platform for early warning through the 4G communication module.

[0069] The present invention also provides a fire detection and early warning system based on improved YOLOv8, including:

[0070] A fire dataset establishment and processing module for establishing a fire dataset, annotating the data in the fire dataset, and dividing it into a training set, a validation set, and a test set;

[0071] A YOLOv8 network improvement module for improving the original YOLOv8 network to obtain an improved YOLOv8 network;

[0072] A YOLOv8 network training module for respectively inputting the training set into the original YOLOv8 network and the improved YOLOv8 network for training; using the validation set to evaluate and adjust the model performance obtained during the two training processes to optimize the model parameters; respectively inputting the test set into the models obtained by training the original YOLOv8 network and the improved YOLOv8 network after being evaluated and adjusted by the validation set for testing to determine the optimal model;

[0073] An optimal model deployment module for deploying the optimal model to the edge AI development board of the ground station to obtain the edge AI development board of the ground station with the deployed optimal model;

[0074] The ground station consists of an edge AI development board, a 5.8G FPV video transmission receiving module, and a fire alarm module;

[0075] A fire detection module for obtaining the real-time video stream of the drone through the drone camera and transmitting it to the edge AI development board of the ground station with the deployed optimal model for fire detection;

[0076] The drone is equipped with a camera, a 5.8G FPV video transmission sending module, a first LoRa wireless communication module, and a GPS module;

[0077] A fire early warning module for using the edge AI development board of the ground station with the deployed optimal model in step S500 to detect a fire signal and issue an early warning through the fire alarm module;

[0078] The fire alarm module consists of an STM32C8T6 development board, a second LoRa wireless communication module, and a 4G communication module.

[0079] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0080] 1. The present invention improves the neck structure of the original YOLOv8 network by designing the SMD-neck structure, effectively reducing the number of model parameters while significantly enhancing the model's representation ability for multi-scale target features and remarkably improving the detection accuracy of small-scale flame targets.

[0081] 2. The present invention replaces the convolutional layer of the backbone network of the original YOLOv8 network by introducing the SPD-Conv module, effectively strengthening local details while significantly reducing the number of model parameters and computational complexity, and reducing the performance requirements of the model for edge computing devices.

[0082] 3. The present invention constructs a multi-modal disaster situation transmission system with a heterogeneous communication architecture, integrating 5.8GHz FPV video transmission, LoRaWAN, and 4G CAT1 links to achieve monitoring coverage of 10 km² in mountainous areas without network coverage.

[0083] 4. The present invention obtains the longitude and latitude information of the fire in real time through the GPS module, and synchronously uploads the location data and fire images to the cloud platform in combination with the 4G communication module, supporting rapid positioning and emergency resource scheduling, and significantly shortening the emergency response time.

[0084] In summary, the present invention improves the YOLOv8 network, enhances the detection accuracy of small-scale fires and significantly improves the deployment efficiency of the algorithm on edge devices. At the same time, a small target fire detection and air-ground collaborative real-time early warning system is established, realizing the efficient collaboration between drones and ground devices, and providing a high-precision and low-latency intelligent solution for early fire warning. BRIEF DESCRIPTION OF THE DRAWINGS

[0085] Figure 1 is a structural block diagram of the fire detection system of the present invention.

[0086] Figure 2 is a schematic structural diagram of the original YOLOv8 network.

[0087] Figure 3 is a schematic structural diagram of the improved YOLOv8 network of the present invention.

[0088] Figure 4 is a schematic diagram of the Slim-neck structure.

[0089] Figure 5 is a schematic diagram of the GSConv module structure.

[0090] Figure 6(a) is a schematic diagram of the structure of the VoV-GSCSP module, and Figure 6(b) is a schematic diagram of the structure of the GS bottleneck module.

[0091] Figure 7 is a schematic diagram of the structure of the MLCA hybrid local channel attention.

[0092] Figure 8 It is a schematic diagram of the dynamic upsampling process of the DySample upsampler.

[0093] Figure 9(a) is the structural diagram of the static sampling point generator of the DySample upsampler, and Figure 9(b) is the structural diagram of the dynamic sampling point generator of the DySample upsampler.

[0094] Figure 10 It is a schematic diagram of the structure of the SMD-neck structure of the present invention.

[0095] Figure 11 It is a schematic diagram of the structure of the SPD-Conv module of the present invention.

[0096] Figure 12(a) is the detection result diagram of the original YOLOv8 network for urban fires, and Figure 12(b) is the detection result diagram of the improved YOLOv8 network for urban fires.

[0097] Figure 13(a) is the detection result diagram of the original YOLOv8 network for forest fires, and Figure 13(b) is the detection result diagram of the improved YOLOv8 network for forest fires.

[0098] Figure 14 It is a schematic diagram of the operation of the fire alarm module. Detailed implementation manners

[0099] In order to enable those skilled in the art to better understand the solutions of the embodiments of the present invention, the following further detailed description of the embodiments of the present invention will be given in conjunction with the accompanying drawings and implementation manners.

[0100] As Figure 1 shown, a fire detection and early warning method based on improved YOLOv8 includes the following steps:

[0101] S100: Establish a fire dataset, annotate the data in the fire dataset, and divide it into a training set, a validation set, and a test set;

[0102] The specific method of step S100 includes:

[0103] S101: Collect flame and smoke images in different weather, different light, and different regions to construct a fire dataset;

[0104] S102: Screen the data in the fire dataset constructed in step S101 and delete low-quality data;

[0105] S103: Annotate the screened fire dataset;

[0106] S104: To prevent overfitting, the fire dataset labeled in step 103 is randomly divided into a training set, a validation set, and a test set according to the ratio of 7:2:1.

[0107] In view of the lack of the currently public remote sensing and UAV fire datasets, the present invention constructs a dedicated image dataset for the fire detection task. The original fire image data is obtained through a multi-source acquisition method, and a screening mechanism based on image sharpness evaluation and content relevance analysis is adopted to eliminate low-quality and irrelevant samples. Finally, 6,744 effective samples are retained. The labelimg annotation tool is used to perform double-category (fire / smoke) bounding box annotation to ensure annotation consistency. Finally, the dataset is divided into a training set (4,721 images), a validation set (1,349 images), and a test set (674 images) according to the ratio of 7:2:1 to ensure the balance of category distribution. The establishment of this dataset provides reliable data support for the feature learning of fire detection algorithms.

[0108] S200: Improve the original YOLOv8 network to obtain the improved YOLOv8 network;

[0109] The specific method of step S200 includes:

[0110] The YOLOv8 algorithm is one of the latest algorithms in the YOLO series. Figure 2 It is a schematic diagram of the network structure of the YOLOv8 algorithm, which can be roughly divided into three parts: the backbone network, the neck structure, and the detection head. Although the YOLOv8 algorithm performs well in object detection tasks, its detection ability for small-scale fire targets still has limitations, and its deployment efficiency on edge devices is relatively low.

[0111] S2001: Introduce the SPD-Conv module to replace the second, third, fourth, and fifth convolutional layers in the backbone network of the original YOLOv8 network;

[0112] In step S2001, introduce the SPD-Conv module to replace the second, third, fourth, and fifth convolutional layers in the backbone network of the original YOLOv8 network; as Figure 2As shown in the figure, the network of the original YOLOv8 algorithm includes a backbone network, a neck structure, and a detection head. The working process of the backbone network is as follows: The input feature map passes through the first convolutional layer (Conv), the second convolutional layer (Conv), the first C2f module, the third convolutional layer (Conv), and the second C2f module in sequence, and then the third feature map (P3) is obtained. The third feature map (P3) is input into the neck structure. The third feature map (P3) then passes through the fourth convolutional layer (Conv) and the third C2f module in sequence, and then the fourth feature map (P4) is obtained. The obtained fourth feature map (P4) is input into the neck structure. The fourth feature map (P4) then passes through the fifth convolutional layer (Conv), the fourth C2f module, and the SPPF module in sequence, and then the fifth feature map (P5) is obtained. The obtained fifth feature map (P5) is input into the neck structure. After feature fusion in the neck structure, the fused features are input into the detection head for detection work.

[0113] As Figure 3 shown in the figure, the working process of the backbone network of the improved YOLOv8 network includes:

[0114] The input feature map passes through the first convolutional layer (Conv), the second SPD-Conv module, the first C2f module, the third SPD-Conv module, and the second C2f module in sequence, and then the third feature map (P3) is obtained. The third feature map (P3) is input into the neck structure. The third feature map (P3) then passes through the fourth SPD-Conv module and the third C2f module in sequence, and then the fourth feature map (P4) is obtained. The obtained fourth feature map (P4) is input into the neck structure. The fourth feature map (P4) then passes through the fifth SPD-Conv module, the fourth C2f module, and the SPPF module in sequence, and then the fifth feature map (P5) is obtained. The obtained fifth feature map (P5) is input into the neck structure.

[0115] As Figure 11As shown, the workflow of the SPD-Conv module (Yang Z, Wu Q, Zhang F, et al. A New Semantic Segmentation Method for Remote Sensing Images Integrating Coordinate Attention and SPD-Conv[J]. Symmetry (20738994), 2023, 15(5). DOI: 10.3390 / sym15051037.) includes: an input feature map with the number of channels , height and width passes through the Space-to-Depth (SPD) layer, which rearranges the spatial blocks of pixels into the depth dimension, increasing the number of channels to , and reducing the spatial dimension to of the original. Then, different channel groups are merged in the depth dimension, and the merged feature map may be added to other processed feature maps (not shown in the figure). Finally, a non-strided convolutional layer performs a convolution with a stride of 1 on the feature map, reducing the channel dimension to , and the spatial resolution remains of the original, obtaining the output feature map.

[0116] Through the SPD layer and the non-strided convolutional layer, the SPD-Conv module can capture and retain the fine information that is often lost when processing small objects and low-resolution images to the greatest extent.

[0117] S2002: Construct an SMD-neck structure to replace the neck structure (Neck) of the original YOLOv8 network; specifically including:

[0118] Introduce the Slim-neck structure. As Figure 4 shown, the Slim-neck structure (Li H, Li J, Wei H, et al. Slim-neck by GSConv: a lightweight-design for real-time detector architectures[J]. Journal of Real-Time Image Processing, 2024, 21(3). DOI: 10.1007 / s11554-024-01436-6.) is a structure designed to optimize the neck structure in convolutional neural networks. It constructs an efficient neural network "neck" through the VoV-GSCSP module and the GSConv module.

[0119] Introduce the MLCA hybrid local channel attention in the Slim-neck structure, that is, add the MLCA hybrid local channel attention after the second, third, and fourth layers of the VoV-GSCSP modules in the Slim-neck structure respectively to obtain the SM-neck structure;

[0120] Introduce the DySample upsampler to replace the Upsample upsampler in the SM-neck structure, that is, replace the first and second layer Upsample upsamplers in the SM-neck structure with the DySample upsampler to obtain the SMD-neck structure;

[0121] Compared with the Slim-neck structure, the SMD-neck improves the feature expression ability of the structure while maintaining its efficient performance, especially in the expression of small target features.

[0122] Replace the neck structure (Neck) of the original YOLOv8 network with the obtained SMD-neck structure.

[0123] In the step S2002, construct the SMD-neck structure to replace the neck structure (Neck) of the original YOLOv8 network, where, as Figure 10As shown in the figure, the working process of the SMD-neck structure is as follows: After the fifth-layer feature map (P5) output by the backbone network of the improved YOLOv8 network in step S2002 passes through the first-layer DySample upsampler, it is concatenated with the fourth-layer feature map (P4) output by the backbone network in the channel dimension through the first-layer feature fusion layer (Concat) to obtain a new feature map PA; The feature map PA passes through the first-layer VoV-GSCSP module and the second-layer DySample upsampler in sequence, and then is concatenated with the third-layer feature map (P3) output by the backbone network in the channel dimension through the second-layer feature fusion layer (Concat) to obtain a new feature map PB; The feature map PB passes through the second-layer VoV-GSCSP module and the first-layer MLCA hybrid local channel attention in sequence to obtain a feature map PC, and the feature map PC is input to head-1 of the detection head; After the feature map PC passes through the first-layer GSConv module, it is concatenated with the feature map output by the first-layer VoV-GSCSP module in the channel dimension through the third-layer feature fusion layer (Concat). The output feature map passes through the third-layer VoV-GSCSP module and the second-layer MLCA hybrid local channel attention in sequence to obtain a feature map PD, and the feature map PD is input to head-2 of the detection head; After the feature map PD passes through the second-layer GSConv module, it is concatenated with the fifth-layer feature map (P5) in the channel dimension through the fourth-layer feature fusion layer (Concat). The output feature map passes through the fourth-layer VoV-GSCSP module and the third-layer MLCA hybrid local channel attention in sequence to obtain a feature map PE, and the feature map PE is input to head-3 of the detection head;

[0124] By introducing the MLCA hybrid local channel attention and the DySample upsampler, while maintaining the light weight of the Slim-neck structure, the SMD-Neck structure significantly improves the efficiency of feature extraction and upsampling, thereby improving the overall detection performance, enhancing the adaptability to complex scenes, being able to better handle the detection problem of multi-scale targets, and especially improving the detection accuracy of small target objects.

[0125] As Figure 5 shown, the working process of the GSConv module is as follows: The input feature map with the number of channels passes through the convolutional layer (Conv) to obtain a feature map with the number of channels, and then passes through the depthwise separable convolutional layer (DWConv) and is concatenated with the feature map with the number of channels output by the convolutional layer (Conv) in the feature fusion layer (Concat). After the concatenated feature map undergoes a shuffle operation, a feature map with The output feature map of the number of channels. The GSConv module can accelerate the prediction calculation of images in a convolutional neural network. Compared with the channel-dense convolution (SC) that maximally preserves the cross-channel implicit correlation and the complete cut-off mechanism of the channel-sparse convolution (DSC), the GSConv module significantly reduces the time complexity while maintaining the feature expressiveness by partially preserving the key channel interaction relationships. The time complexity is usually defined by the floating-point operations (FLOPs). Therefore, the time complexities of the channel-dense convolution (SC), the channel-sparse convolution (DSC), and the GSConv module are respectively:

[0126]

[0127]

[0128]

[0129] Among them, is the width of the output feature map, is the height, is the size of the convolution kernel, is the number of channels of each convolution kernel and also the number of channels of the input feature map, is the number of channels of the output feature map; the computational cost of the GSConv module is about 50% of that of the channel-dense convolution (SC), but its contribution to the model learning ability is equivalent to that of the channel-dense convolution (SC).

[0130] As shown in Fig. 6(a), the working process of the VoV-GSCSP module includes: the input feature map passes through the first convolutional layer (Conv) and the GS bottleneck module in sequence, and then is concatenated with the output feature map of the first convolutional layer (Conv) at the feature fusion layer (Concat). The concatenated feature map passes through the second convolutional layer (Conv) to obtain the output feature map. The VoV-GSCSP module can improve the feature utilization efficiency and network characteristics, and perform effective information fusion between the feature maps at different stages. The design of the VoV-GSCSP module is based on the GS bottleneck module.

[0131] As shown in Fig. 6(b), the working process of the GS bottleneck module includes: the input feature map passes through the convolutional layer (Conv) to obtain the feature map Pa; the input feature map passes through two GSConv modules in sequence to obtain the feature map Pb; after the addition operation of the feature map Pa and the feature map Pb, the output feature map is obtained. The GS bottleneck module is based on the GSConv module and is an enhancement module used to improve the non-linear expression of features and the reuse of information, and improve the learning ability of the model by stacking the GSConv modules.

[0132] As Figure 7 shown, the working process of the MLCA (Mixed Local Channel Attention, Wan Dahang, Lu Rongsheng, Shen Siyuan, et al. Mixed local channel attention for object detection [J]. 2023.) includes: The input feature map is first subjected to local average pooling (LAP) and global average pooling (GAP) processes in sequence. For the feature map after local average pooling (LAP) processing, it is successively reshaped, 1D-convolved (Conv1d), and reshaped to obtain the feature map Pc; for the feature map after global average pooling (GAP) processing, after 1D-convolution (Conv1d), an unpooling (UNAP) operation is performed, and then it is added to the feature map Pc. The feature map output from the addition operation is subjected to an unpooling (UNAP) operation and then multiplied by the original input feature map to obtain the output feature map.

[0133] The MLCA (Mixed Local Channel Attention) performs local average pooling and global average pooling on the input feature map respectively. After compressing the channels and maintaining the spatial dimensions through 1D convolution, the local features are fused with the original input through multiplication to strengthen the key regions, and the global features are combined with the local results through addition to inject global statistics. Finally, the original size is restored through unpooling to achieve the collaborative optimization of local fine features and global semantic information.

[0134] The DySample upsampler for the input feature Its upsampling process is expressed as:

[0135]

[0136] where is the sampling set generated by the sampling point generator, is the output; the DySample upsampler generates the sampling set through two types of sampling point generators, static and dynamic;

[0137] As Figure 8 shown, the working process of the DySample upsampler includes: The input feature map creates the sampling set through the sampling point generator, and the sampling set and the input feature map are operated through the grid_sample function for resampling to obtain the upsampled feature map;

[0138] The sampling point generator includes a static range factor and a dynamic range factor; where:

[0139] As shown in Figure 9(a), the workflow of the static range factor: The input feature map passes through a linear layer and is combined with a fixed range factor (0.25), and then the offset is generated through the pixel shuffle technique ( ), and the offset ( ) is added to the original network position ( ) to obtain the sampling set ( );

[0140] As shown in Figure 9(b), the workflow of the dynamic range factor: The input feature map passes through a linear layer and is combined with the dynamic range factor ( ), the obtained output is multiplied by the feature map passing through the linear layer, and then the offset is generated through the pixel shuffle technique ( ), and the offset ( ) is added to the original network position ( ) to obtain the sampling set ( ), where represents the Sigmoid function, which is used to generate the range factor.

[0141] The whole process can be defined as:

[0142]

[0143]

[0144] Replace the neck structure (Neck) of the original YOLOv8 network with the SMD-neck structure to obtain the improved YOLOv8 network. Figure 3 is the network structure of the improved YOLOv8 algorithm. The SPD-Conv module is introduced to replace the second, third, fourth, and fifth convolutional layers (Conv) in the backbone network, and then the MLCA hybrid local channel attention mechanism and the DySample upsampler are introduced to improve the Slim-neck structure. The SMD-neck structure is designed to replace the neck structure (Neck) of the original YOLOv8 network to obtain the network structure of the improved YOLOv8 algorithm;

[0145] S300: Input the training sets divided in step S100 into the original YOLOv8 network and the improved YOLOv8 network in step S200 for training respectively; use the validation set divided in step S100 to evaluate and adjust the model performance obtained during the two training processes to optimize the model parameters; input the test sets divided in step S100 into the models obtained by training the original YOLOv8 network evaluated and adjusted by the validation set and the models obtained by training the improved YOLOv8 network for testing to determine the optimal model;

[0146] Step S300 includes the following steps:

[0147] S301: Input the training set divided in step S100 into the original YOLOv8 network for training;

[0148] S302: Input the training set divided in step S100 into the improved YOLOv8 network in step S200 for training;

[0149] S303: Compare the training results of steps S301 and S302 to obtain the optimal model; The main model parameters to be compared are the number of parameters (Parameters), the number of floating-point operations (GFLOPs), and the mean average precision (mAP). Among them, the number of parameters (Parameters) and the number of floating-point operations (GFLOPs) are directly related to the complexity, computational requirements, and storage requirements of the model, while the mean average precision (mAP) is usually used as a comprehensive indicator to evaluate the performance of the model.

[0150] The specific comparison method is: Introduce the APG coefficient to evaluate the comprehensive performance of the model. The larger this index, the better the performance; The calculation formula of the APG coefficient is:

[0151]

[0152] where, , , are weight coefficients used to measure the importance of the three indicators, is the number of parameters of the model obtained by training the original YOLOv8 network, is the number of floating-point operations of the model obtained by training the original YOLOv8 network.

[0153] For the original YOLOv8 network and the improved YOLOv8 network, the present invention conducted the following experiments:

[0154] 1) The present invention wrote the algorithm based on the Pytorch framework, and the experiment was conducted in the environment shown in Table 1.

[0155] Table 1 Experimental environment

[0156]

[0157] 2) Fire dataset

[0158] Use the dataset divided in S100 for training.

[0159] 3) Evaluation indicators

[0160] The present invention sets the intersection over union (IoU) to 0.5, that is, when the IoU between the predict box and the ground-truth is greater than 0.5, it indicates a successful prediction. mAP, Parameters, and GFLOPs are used as the model evaluation criteria. mAP is the average of the average precisions of all classes in the dataset. In object detection tasks, mAP is usually used as a comprehensive indicator to evaluate the performance of the model. Parameters and GFLOPs are directly related to the complexity, computational requirements, and storage requirements of the model. In scenarios with computational resource limitations, it may be necessary to trade off between the accuracy and computational efficiency of the model.

[0161] 4) Experimental results and analysis

[0162] 4.1) Model hyperparameter settings are shown in Table 2:

[0163] Table 2 Model hyperparameters

[0164]

[0165] 4.2) Experimental results and analysis

[0166] Table 3 Experimental results of different improvement strategies

[0167]

[0168] Among them, the calculation formula for the APG coefficient is:

[0169]

[0170] Among them, Take 1, 、 Take 0.2, Take 3011238, Take 8.20.

[0171] Figure 12(a) shows the training results of the original YOLOv8 network on this dataset, and Figure 12(b) shows the training results of the improved YOLOv8 network.

[0172] The present invention realizes the collaborative optimization of the YOLOv8 network in terms of accuracy and efficiency through modular improvement, and trains the optimal model. Experimental data shows that:

[0173] The SMD-neck structure achieves a +3% mAP improvement compared to the model trained by the original YOLOv8 network, reducing the number of parameters by 200,000; compared to the Slim-neck structure, it achieves a +2.4% mAP improvement under the condition of reducing a small number of parameters.

[0174] The SPD-Conv module uses spatial-to-depth transformation to achieve a 217,808 reduction in parameters, a 0.5 reduction in GFLOPs, and a +2.2% mAP breakthrough, verifying the effectiveness of cross-scale feature preservation;

[0175] The optimal model obtained by jointly deploying and training the SMD-neck structure and the SPD-Conv module has a significant compression effect of reducing 417K in parameters and 1.3 in GFLOPs, while achieving an accuracy breakthrough of +3.1% mAP. This shows that the improved YOLOv8 network of the present invention can greatly improve the detection accuracy while significantly reducing the demand for device hardware storage and the computational burden.

[0176] In order to demonstrate that the improved YOLOv8 network of the present invention can ensure detection accuracy while reducing the number of parameters and floating-point calculations, based on the fire dataset in step S100, the YOLOv5s algorithm, YOLOv6 algorithm, YOLOv7-tiny algorithm, YOLOv8 algorithm, YOLOv8-p6 algorithm and the improved YOLOv8 algorithm of the present invention are respectively introduced for training, and the obtained models are compared. The experimental data are as follows:

[0177] Table 4 Experimental results of the improved YOLOv8 algorithm and other YOLO algorithms

[0178]

[0179] As shown in the table, the mAP of the model obtained by training the improved YOLOv8 network of the present invention is higher than that of the models obtained by the current mainstream single-stage object detection algorithms, and the parameters and GFLOPs are lower than those of the models obtained by the current mainstream single-stage object detection algorithms.

[0180] To demonstrate the performance of the improved YOLOv8 network for fire detection, the present invention selects fire pictures to test the models obtained by training the improved YOLOv8 network and the original YOLOv8 network. The present invention selects urban fire images from an aerial perspective. Figure 12(a) is the detection result diagram of the original YOLOv8 network for urban fires, and Figure 12(b) is the detection result diagram of the improved YOLOv8 network for urban fires. By comparison, it can be found that in the urban fire detection task from an aerial perspective, for the flames and smoke that can be detected by the model obtained by training the original YOLOv8 network, the model obtained by training the improved YOLOv8 network of the present invention has higher detection accuracy; for the flames and smoke that cannot be detected by the model obtained by training the original YOLOv8 network, the model obtained by training the improved YOLOv8 network of the present invention can detect them. The present invention selects drone-captured forest fire images. Figure 13(a) is the detection result diagram of the original YOLOv8 network for forest fires, and Figure 13(b) is the detection result diagram of the improved YOLOv8 network for forest fires. By comparison, it can be found that in the forest fire detection task from the drone's downward perspective, the model obtained by training the improved YOLOv8 network of the present invention has significantly better detection effects for both flames and smoke than the model obtained by training the original YOLOv8 network, especially for small-target flames.

[0181] Therefore, through comparison, the present invention obtains the optimal model obtained by training the improved YOLOv8 network.

[0182] S400: Deploy the optimal model obtained in step S300 to the Orange Pi AIpro development board of the ground station to obtain the Orange Pi AIpro development board of the ground station with the optimal model deployed; the Orange Pi AIpro development board is one of the optional solutions for edge AI development boards.

[0183] The Orange Pi AIpro development board is equipped with the ubuntu operating system, adopts the Ascend AI technology route, has rich interfaces and strong scalability, and provides 8 TOPS of powerful computing power.

[0184] S500: Obtain the real-time video stream of the drone through the drone camera and transmit it to the edge AI development board of the ground station with the optimal model deployed in step S400 for fire detection;

[0185] The drone is equipped with a wide-angle optical sensor to capture real-time images. The 5.8G FPV video transmission module is used to transmit the real-time video stream to the 5.8G FPV video reception module connected to the Orange Pi AIpro development board of the ground station, and then transmitted to the improved YOLOv8 fire detection model on the Orange Pi AIpro development board for fire detection. The measured maximum transmission distance can reach 8 kilometers and the end-to-end transmission delay is less than 0.1s, which can meet the needs of beyond-line-of-sight operations in mountainous areas without public network coverage.

[0186] S600: When the edge AI development board of the optimal model deployed at the ground station in step S500 detects a fire signal, it issues a warning through the fire alarm module.

[0187] Step S600 includes the following steps:

[0188] S601: The Orange pi AIpro development board of the optimal model deployed at the ground station detects a fire signal and conducts serial communication through the USART port of the Orange pi AIpro development board with the STM32C8T6 development board of the fire alarm module;

[0189] S602: After receiving the instruction, the STM32C8T6 development board communicates with the first LoRa wireless communication module carried on the drone through the second LoRa wireless communication module of the fire alarm module;

[0190] S603: After receiving the instruction, the first LoRa wireless communication module on the drone transmits the longitude and latitude information of the GPS module back to the STM32C8T6 development board of the fire alarm module;

[0191] S604: The STM32C8T6 development board of the fire alarm module uploads the longitude and latitude information and the fire image to the fire monitoring cloud platform through the 4G communication module for warning.

[0192] After the improved YOLOv8 fire detection model recognizes a fire signal, it sends an alarm message and the intercepted fire picture through the USART serial communication interface of the Orange Pi AIpro development board to the fire alarm module.

[0193] Figure 14 It is a schematic structural diagram of the fire alarm module, which consists of an STM32C8T6 development board, a second LoRa wireless communication module, and a 4G communication module.

[0194] After the STM32C8T6 development board receives the alarm information through the USART port, it conducts point-to-point communication with the LoRa wireless communication module carried on the drone through the LoRa wireless communication module to obtain the drone's GPS information. The measured communication distance can reach 10 kilometers, and the power consumption is extremely low. The STM32C8T6 development board uploads the GPS information and fire pictures to the fire detection cloud platform built based on Alibaba Cloud through the 4G communication module for fire warning.

[0195] Through the heterogeneous networking of LPWAN and cellular networks, this architecture achieves monitoring coverage of unmanned areas at the 10 km² level while ensuring a sub-second response speed, showing a huge improvement compared to traditional solutions.

[0196] The present invention also provides a fire detection and warning system based on improved YOLOv8, including:

[0197] A fire dataset establishment and processing module for establishing a fire dataset in step S100, annotating the data in the fire dataset, and dividing it into a training set, a validation set, and a test set;

[0198] A YOLOv8 network improvement module for improving the original YOLOv8 network in step S200 to obtain an improved YOLOv8 network;

[0199] A YOLOv8 network training module for inputting the training set divided in step S100 into the original YOLOv8 network and the improved YOLOv8 network in step S200 for training respectively in step S300; evaluating and adjusting the model performance obtained during the two training processes using the validation set divided in step S100 to optimize the model parameters; inputting the test set divided in step S100 into the models obtained by training the original YOLOv8 network and the improved YOLOv8 network after being evaluated and adjusted by the validation set for testing to determine the optimal model;

[0200] An optimal model deployment module for deploying the optimal model obtained in step S300 to the edge AI development board of the ground station in step S400 to obtain the edge AI development board of the ground station with the deployed optimal model;

[0201] The ground station consists of an edge AI development board, a 5.8G FPV video transmission receiving module, and a fire alarm module;

[0202] A fire detection module for obtaining the real-time video stream of the drone through the drone camera in step S500 and transmitting it to the edge AI development board of the ground station with the deployed optimal model in step S400 for fire detection;

[0203] The drone is equipped with a camera, a 5.8G FPV video transmission module, a first LoRa wireless communication module, and a GPS module;

[0204] A fire warning module, which is used to implement the detection of a fire signal by the edge AI development board of the optimal deployment model of the ground station in step S500 in step S600, and issue a warning through the fire alarm module;

[0205] The fire alarm module is composed of an STM32C8T6 development board, a second LoRa wireless communication module, and a 4G communication module.

[0206] Aiming at the problems of insufficient detection accuracy of small-scale ignition points in fires, the detection model not being lightweight enough, and high requirements for device performance in existing solutions, the present invention provides a fire detection and warning method and system based on improved YOLOv8. By introducing the MLCA hybrid local channel attention and DySample upsampler to optimize the Slim-neck structure, the SMD-neck structure is designed, which improves the detection accuracy of small-scale targets while reducing the number of model parameters; introducing the SPD-Conv module, by reducing information loss and improving the accuracy of feature extraction, the processing ability of the model for small target objects and low-resolution images is optimized, and the computational complexity is reduced.

[0207] The key points and protected points of the present invention are:

[0208] The present invention takes into account various influencing factors such as the detection of small-scale targets in drone fire detection and the performance limitations of detection devices. Based on a complex fire environment, a rich fire dataset is constructed, and a lightweight small target detection YOLOv8 network is designed. This network can be applied to different fire detection environments, can detect small-scale ignition points with high accuracy, and has the characteristics of being lightweight, with less performance requirements for devices. On this basis, the present invention constructs a lightweight small target fire detection and air-ground collaborative real-time warning system to achieve real-time monitoring of fires.

Claims

1. A fire detection and early warning method based on improved YOLOv8, characterized in that: The following steps are involved: S100: Establish a fire data set, annotate the data in the fire data set, and divide it into a training set, a validation set, and a test set; S200: improving the original YOLOv8 network to obtain an improved YOLOv8 network; the specific method includes: S2001: Introduce the SPD-Conv module to replace the second, third, fourth and fifth convolutional layers in the backbone network Backbone of the original YOLOv8 network; the original YOLOv8 network includes the backbone network Backbone, the neck structure Neck and the detection head Head; S2002: Construct the SMD-neck structure to replace the neck structure of the original YOLOv8 network; specifically include: Introduce the Slim-neck structure, and introduce the MLCA hybrid local channel attention into the Slim-neck structure, that is, add the MLCA hybrid local channel attention to the second, third, and fourth layers of the VoV-GSCSP module in the Slim-neck structure to obtain the SM-neck structure; The DySample upsampler is introduced to replace the Upsample upsampler in the SM-neck structure, that is, the DySample upsampler replaces the first and second layers of the Upsample upsampler in the SM-neck structure to obtain the SMD-neck structure; The obtained SMD-neck structure replaces the neck structure Neck of the original YOLOv8 network; S300: input the training set divided in step S100 into the original YOLOv8 network and the improved YOLOv8 network after step S200 for training respectively; use the validation set divided in step S100 to evaluate and adjust the model performance obtained in the two training processes to optimize the model parameters; input the test set divided in step S100 into the model obtained by the original YOLOv8 network training and the model obtained by the improved YOLOv8 network training after the validation set evaluation and adjustment respectively for testing to determine the optimal model; S400: deploying the optimal model obtained in step S300 to the edge AI development board of the ground station to obtain the edge AI development board of the ground station where the optimal model is deployed; S500: Obtain the real-time video stream of the drone through the drone camera and transmit it to the edge AI development board where the optimal model is deployed at the ground station in step S400 for fire detection; S600: The edge AI development board of the optimal model deployed at the ground station in step S500 detects a fire signal and issues an early warning through the fire alarm module.

2. A fire detection and early warning method based on improved YOLOv8 according to claim 1, characterized in that: The specific method of step S100 includes: S101: Collect fire and smoke images in different weather conditions, light conditions, and regions to build a fire dataset; S102: Screening the data in the fire data set constructed in step S101 and deleting low-quality data; S103: labeling the filtered fire data set; S104: Randomly divide the fire data set annotated in step 103 into a training set, a validation set and a test set.

3. A fire detection and early warning method based on improved YOLOv8 according to claim 1, characterized in that: In step S2001, the SPD-Conv module is introduced to replace the second, third, fourth and fifth convolutional layers in the backbone network Backbone of the original YOLOv8 network; the workflow of the backbone network Backbone of the improved YOLOv8 network includes: The input feature map passes through the first convolution layer Conv, the second SPD-Conv module, the first C2f module, the third SPD-Conv module, and the second C2f module in sequence to obtain the third-layer feature map P3, and the third-layer feature map P3 is input into the neck structure Neck; the third-layer feature map P3 passes through the fourth SPD-Conv module and the third-layer C2f module in sequence to obtain the fourth-layer feature map P4, and the fourth-layer feature map P4 is input into the neck structure Neck; the fourth-layer feature map P4 passes through the fifth SPD-Conv module, the fourth-layer C2f module, and the SPPF module in sequence to obtain the fifth-layer feature map P5; the fifth-layer feature map P5 is input into the neck structure Neck; In the step S2002, an SMD-neck structure is constructed to replace the neck structure Neck of the original YOLOv8 network, wherein the workflow of the SMD-neck structure includes: the fifth-layer feature map P5 output by the backbone network Backbone of the YOLOv8 network improved by step S2002 passes through the first-layer DySample upsampler, and is then spliced ​​with the fourth-layer feature map P4 output by the backbone network Backbone in the channel dimension through the first-layer feature fusion layer Concat to obtain a new feature map PA; the feature map PA passes through the first-layer VoV-GSCSP module and the second-layer DySample upsampler in sequence, and is then spliced ​​with the third-layer feature map P3 output by the backbone network Backbone in the channel dimension through the second-layer feature fusion layer Concat to obtain a new feature map PB; the feature map PB passes through the second-layer VoV-GSC After the SP module and the first layer of MLCA mixed local channel attention, the feature map PC is obtained, and the feature map PC is input into the head-1 of the detection head; after the feature map PC passes through the first layer of GSConv module, it is concatenated with the feature map output by the first layer of VoV-GSCSP module through the third layer of feature fusion layer Concat in the channel dimension, and the output feature map passes through the third layer of VoV-GSCSP module and the second layer of MLCA mixed local channel attention in turn to obtain the feature map PD, and the feature map PD is input into the head-2 of the detection head; after the feature map PD passes through the second layer of GSConv module, it is concatenated with the fifth layer of feature map P5 in the channel dimension through the fourth layer of feature fusion layer Concat, and the output feature map passes through the fourth layer of VoV-GSCSP module and the third layer of MLCA mixed local channel attention in turn to obtain the feature map PE, and the feature map PE is input into the head-3 of the detection head; The SMD-neck structure replaces the neck structure Neck of the original YOLOv8 network.

4. A fire detection and early warning method based on improved YOLOv8 according to claim 3, characterized in that: In step S2001, the workflow of the SPD-Conv module includes: the input feature map passes through the space-to-depth SPD layer, the spatial blocks of pixels are rearranged to the depth dimension, and then different channel groups are merged in the depth dimension, and finally the output feature map is obtained through the non-strided convolution layer; In step S2002, the workflow of the GSConv module includes: the input feature map with c1 channels is passed through the convolution layer Conv to obtain a feature map with c1 channels. The feature map of the number of channels is then passed through the depth-separable convolutional layer DWConv and the convolutional layer Conv output has The feature maps with the same number of channels are concatenated in the feature fusion layer Concat. After the concatenated feature maps are randomly permuted and shuffled, an output feature map with c2 channels is obtained. The workflow of the VoV-GSCSP module includes: after the input feature map passes through the first convolution layer Conv and the GSBottleneck module in sequence, it is spliced ​​with the output feature map of the first convolution layer Conv in the feature fusion layer Concat, and the spliced ​​feature map passes through the second convolution layer Conv to obtain the output feature map; The workflow of the GS bottleneck module includes: the input feature map passes through the convolution layer Conv to obtain the feature map Pa; the input feature map passes through two layers of GSConv modules continuously to obtain the feature map Pb; the feature map Pa and the feature map Pb are added to obtain the output feature map; The workflow of the MLCA hybrid local channel attention includes: the input feature map is first processed by local average pooling LAP and global average pooling GAP in sequence, and the feature map after the local average pooling LAP processing is sequentially subjected to rearrangement Reshape, 1D convolution Conv1d, and rearrangement Reshape to obtain the feature map Pc; the feature map after the global average pooling GAP processing is subjected to unpooling UNAP operation after 1D convolution Conv1d, and then added to the feature map Pc, and the feature map output by the addition operation is subjected to unpooling UNAP operation and then multiplied with the original input feature map to obtain the output feature map; The upsampling process of the DySample upsampler for the input feature x is expressed as: x' = grid_sample(x,S) Among them, S is the sampling set generated by the sampling point generator, and x' is the output; the DySample upsampler generates the sampling set through static and dynamic sampling point generators; The workflow of the DySample upsampler includes: the input feature map is used to create a sampling set through a sampling point generator, and the sampling set and the input feature map are operated through a grid_sample function to perform resampling to obtain an upsampled feature map.

5. A fire detection and early warning method based on improved YOLOv8 according to claim 4, characterized in that: The sampling point generator includes a static range factor and a dynamic range factor; wherein: The workflow of the static range factor includes: the input feature map passes through the linear layer and is combined with a fixed range factor of 0.25, and then the pixel shuffle technology is used to generate an offset o, which is added to the original network position g to obtain a sample set S; The workflow of the dynamic range factor includes: the input feature map passes through the linear layer and is combined with the dynamic range factor 0.5σ. The output is multiplied with the feature map after the linear layer, and then the offset o is generated through the pixel shuffle technology. The offset o is added to the original network position g to obtain the sampling set S, where σ represents the Sigmoid function, which is used to generate the range factor.

6. A fire detection and early warning method based on improved YOLOv8 according to claim 1, characterized in that: Step S300 includes the following steps: S301: Input the training set divided in step S100 into the original YOLOv8 network for training; S302: Input the training set divided in step S100 into the improved YOLOv8 network in step S200 for training; S303: Compare the training results of step S301 and S302 to obtain the optimal model; the model parameters compared are parameter quantity Parameters, floating point operation number GFLOPs and mean average precision mAP; the specific comparison method is: introduce the APG coefficient to evaluate the comprehensive performance of the model, the larger the APG coefficient, the better the performance; the calculation formula of the APG coefficient is: Among them, α, β, and γ are weight coefficients used to measure the importance of the three indicators, parameters1 is the parameter quantity of the model obtained by training the original YOLOv8 network, and GFLOPs1 is the number of floating-point operations of the model obtained by training the original YOLOv8 network.

7. A fire detection and early warning method based on improved YOLOv8 according to claim 1, characterized in that: Step S500 includes the following steps: S501: The drone is equipped with a camera, which transmits the captured video stream to the 5.8G FPV image transmission receiving module in real time through the 5.8G FPV image transmission sending module; S502: The edge AI development board of the optimal model deployed in the ground station in step S400 receives the real-time video stream of the drone through the 5.8G FPV image transmission receiving module; S503: Use the edge AI development board of the optimal model deployed at the ground station in step S400 to detect the video stream received in step S502.

8. The fire detection and early warning method based on improved YOLOv8 according to claim 1, characterized in that: The step S600 includes the following steps: S601: The edge AI development board of the optimal deployment model of the ground station detects a fire signal and communicates with the STM32C8T6 development board of the fire alarm module through the USART port of the edge AI development board; S602: The STM32C8T6 development board receives the command and communicates with the first LoRa wireless communication module on the drone through the second LoRa wireless communication module of the fire alarm module; S603: The first LoRa wireless communication module on the drone receives the command and transmits the longitude and latitude information of the GPS module back to the STM32C8T6 development board of the fire alarm module; S604: The STM32C8T6 development board of the fire alarm module uploads the latitude and longitude information and fire images to the fire monitoring cloud platform for early warning through the 4G communication module.

9. A fire detection and early warning system based on improved YOLOv8 based on the method according to any one of claims 1 to 8, characterized in that: include: The fire data set establishment and processing module is used to establish the fire data set, annotate the data in the fire data set, and divide it into training set, verification set and test set; YOLOv8 network improvement module, used to improve the original YOLOv8 network to obtain an improved YOLOv8 network; The YOLOv8 network training module is used to input the training set into the original YOLOv8 network and the improved YOLOv8 network for training; the validation set is used to evaluate and adjust the model performance obtained during the two training processes to optimize the model parameters; the test set is input into the model obtained by the original YOLOv8 network training and the model obtained by the improved YOLOv8 network training after the validation set evaluation and adjustment respectively for testing to determine the optimal model; An optimal model deployment module, used to deploy the optimal model to the edge AI development board of the ground station, and obtain the edge AI development board of the ground station where the optimal model is deployed; The ground station consists of an edge AI development board, a 5.8G FPV image transmission receiving module, and a fire alarm module; The fire detection module is used to obtain the real-time video stream of the drone through the drone camera and transmit it to the edge AI development board deployed with the optimal model at the ground station for fire detection; The drone is equipped with a camera, a 5.8G FPV image transmission module, a first LoRa wireless communication module and a GPS module; A fire warning module is used to detect a fire signal using the edge AI development board of the optimal model deployed by the ground station in step S500, and issue an early warning through a fire alarm module; The fire alarm module is composed of an STM32C8T6 development board, a second LoRa wireless communication module and a 4G communication module.

Citation Information

Patent Citations

  • Unmanned aerial vehicle forest fire detection system

    CN114037910A

  • Fire detection method and system based on unmanned aerial vehicle edge calculation

    CN116206223A

  • Fire detection method based on deep learning, unmanned aerial vehicle and storage medium

    CN116229296A

  • Fire detection method based on PD-YOLO

    CN117333753A

  • Fire inspection early warning method based on unmanned aerial vehicle and improved YOLOv8 target detection algorithm

    CN117409191A