Smoke and fire identification and positioning method and device in unmanned aerial vehicle inspection

By carrying high-definition cameras and infrared sensors on the drone, combined with image processing and deep learning algorithms, dynamic transmission and firework monitoring systems are designed, and the problems of limited coverage and slow response of traditional firework monitoring methods are solved, achieving high-precision firework recognition and fast response in complex environments.

CN120071199APending Publication Date: 2025-05-30ZHEJIANG UNIV OF TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510153877.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional fireworks monitoring methods have limited coverage and slow response speed, especially in complex environments with poor recognition effect.

Method used

The UAV is equipped with high-definition cameras and infrared sensors, combined with image processing and deep learning algorithms, and designed dynamic transmission and firework monitoring systems, and uses CO-DETR models and timing networks to achieve real-time monitoring, precise positioning and rapid response.

Benefits of technology

It improves the accuracy and response speed of firework monitoring, can accurately identify firework phenomena in complex environments, and reduces the risk of fireworks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071199A_ABST
    Figure CN120071199A_ABST
Patent Text Reader

Abstract

The invention discloses a smoke and fire identification and positioning method and device in unmanned aerial vehicle inspection, and the method comprises the steps: firstly, building an unmanned aerial vehicle smoke and fire true color and infrared image fusion data set and a time sequence data set under the same coordinates; then, performing data enhancement by adopting a generative adversarial model; thirdly, constructing a fusion image model by using a CO-DETR model; meanwhile, a time sequence model is designed and covers image feature extraction and dimension reduction, time sequence data vectorization processing, attention calculation of sudden change perception and region discrimination; a weighted voting mechanism is adopted, and prediction results are weighted according to the confidence coefficient of each model; video stream receiving and reasoning tasks are decoupled through a multi-thread technology, meanwhile, the network state is intercepted, and transmission of video streams is dynamically adjusted; and finally, iteratively optimizing the model by adopting an incremental learning mode. A dynamic transmission and smoke and fire monitoring system is designed by combining the autonomous flight capability of the unmanned aerial vehicle and the multi-sensor technology, and meanwhile, real-time monitoring, accurate positioning and quick response of the smoke and fire phenomenon are realized by using a dynamic open source model CO-DETR and a sequential network designed according to the characteristic that the unmanned aerial vehicle has high repeatability in daily cruising.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned aerial vehicle (UAV) applications, and particularly to a method and device for smoke and fire recognition and positioning in UAV patrol inspection, belonging to the field of intelligent patrol inspection and monitoring systems. Specifically, it involves the comprehensive application of image processing, deep learning algorithms, multi-sensor fusion technology, real-time communication, and a combined decision-making module for consecutive frames. Background Art

[0002] With the rapid development of UAV technology, more and more fields have started to use UAVs for environmental monitoring and disaster warning. Traditional smoke and fire monitoring methods, such as ground monitoring and satellite remote sensing, have problems such as limited coverage and slow response speed. Especially in complex or inaccessible areas (such as mountains, forests, and high-rise buildings in cities), the effect is poor. Therefore, there is an urgent need for a more efficient and accurate smoke and fire monitoring system.

[0003] UAVs have a highly flexible flight ability, can cover a wide area, and are equipped with high-definition cameras and infrared sensors for real-time data collection. Combining image processing and deep learning algorithms can effectively identify smoke and fire phenomena. In addition, multi-sensor fusion technology improves the recognition accuracy and reliability of the system in different environments.

[0004] Although some existing UAV smoke and fire monitoring systems have achieved preliminary applications, their response time and monitoring accuracy still have certain limitations. Therefore, there is an urgent need for an innovative UAV smoke and fire monitoring solution to improve the early warning ability and response speed and reduce the risk of smoke and fire disasters. Summary of the Invention

[0005] The present invention aims to overcome the above-mentioned drawbacks of the prior art and provides a method and device for smoke and fire recognition and positioning in UAV patrol inspection, aiming to solve the problems of limited coverage, slow response speed, and poor recognition effect in complex environments of traditional smoke and fire monitoring methods. The present invention combines the autonomous flight ability of UAVs and multi-sensor technology to design a dynamic transmission and smoke and fire monitoring system. At the same time, the dynamic open-source model CO-DETR and a temporal network designed according to the high repeatability of the daily UAV cruise are used to achieve real-time monitoring, precise positioning, and rapid response to smoke and fire phenomena.

[0006] The method of the present invention includes the following technical modules:

[0007] UAV platform: Equipped with devices such as high-definition cameras and infrared sensors to collect real-time environmental data.

[0008] Data Processing and Analysis Module: It uses image processing technology and deep learning algorithms to process images and sensor data, achieving high-precision recognition of fire and smoke targets. It includes a model framework that effectively handles the domain differences between "historical no fire and smoke" and "current fire and smoke" in time-series data. By introducing a domain discriminator, domain labels, dynamic difference feature representation, and mutation-aware attention mechanism, it can solve the time-series mutation problem that traditional models cannot effectively handle and improve the prediction ability of the model at the mutation moment.

[0009] Real-time Communication and Early Warning Module: It transmits the detection results and location information to the ground control center in real time through a wireless network, generating fire and smoke early warning information. At the same time, the system adopts dynamic resolution transmission technology to adjust the resolution of the video stream according to the network condition and data importance. In the case of limited bandwidth, it ensures the balance between the timely transmission of early warning information and image quality. When the network condition is good, it transmits high-definition images to provide more details; while when the bandwidth is limited, it reduces the image resolution to ensure real-time performance and stability, thereby improving the response speed and reliability of the overall system.

[0010] Autonomous Cruise and Path Planning Module: It realizes the inspection of a large area and the precise monitoring of key areas through the autonomous navigation function of the unmanned aerial vehicle.

[0011] Decision-making Module: It identifies potential fire and smoke targets and fire and smoke early warning information by combining the image information of the front and rear frames and time-series analysis methods. This module reduces false alarms and improves the accuracy of fire and smoke target recognition through multi-frame voting on the early warning information of the front and rear frames.

[0012] To achieve the above objectives, the first aspect of the present invention relates to a method for fire and smoke recognition and positioning in unmanned aerial vehicle inspection, specifically including the following steps:

[0013] Step 1, making a data set;

[0014] Step 1.1, front and rear frame data set;

[0015] The video of the unmanned aerial vehicle inspection is frame-extracted, and the infrared image obtained by the infrared sensor is used as the fourth channel and fused into the extracted image frames, thereby generating an inspection image data set. Taking the current frame as a reference, the adjacent front and rear frames are obtained to form the front and rear frame data set D frame :

[0016] D frame ={P 1 , P 2 , …, P n-k , …, P n} (1)

[0017] where P i, where \(i = 1, 2, \ldots, n - k, \ldots, n\) represents the image data of the \((n - k)\)-th frame and its adjacent frames in the video. Each frame image contains RGB channels and an infrared image as the fourth channel.

[0018] Step 1.2, Temporal sequence dataset;

[0019] During the flight of the drone, the system records the GPS coordinates and timestamps of the current position. By associating the historical data at the same position, a temporal sequence dataset \(D\) is created. seq :

[0020] \(D\) seq =\(\{(P\) 1 , \((lon\) 1 , lat\) 1 , alt\) 1 ), \(t\) 1 ), \((P\) 2 , \((lon\) 2 , lat\) 2 , alt\) 1 ), \(t\) 2 ), \(\ldots, (P\) n , \((lon\) n , lat\) n , alt\) n ), \(t\) n )\}\ (2)

[0021] where \(P\) i is the image data of the \(i\)-th frame, \((lon\) i , lat\) i , alt\) i ) are the coordinates when the \(i\)-th frame image is taken, and \(t\) i is the timestamp of the \(i\)-th frame image.

[0022] Step 1.3, Data augmentation;

[0023] The collected images are augmented, including conventional methods such as image flipping, translation, scaling, and adding noise. At the same time, the Mosaic data augmentation technique is introduced to generate more training samples.

[0024] To further improve the adaptability of the model to the domain differences between "smokeless" and "smoky", a generative adversarial network is used to generate data samples with different domain characteristics, thereby enhancing the model's adaptive ability to domain differences. The generative adversarial network consists of two main parts: a generator and a discriminator.

[0025] The generator (\(G\)) samples a vector \(z \sim p\) z from the latent space, where \(p_z\) is the Gaussian distribution of the latent space. Then, the generator uses this vector \(z\) as input to generate a sample \(G\) θ(z). By generating samples G θ (z) enables the discriminator D φ to consider these samples as real as possible, and the output D of the discriminator for the generated samples φ (G θ (z)) is close to 1. The generator loss function L G is defined as:

[0026]

[0027] The discriminator D φ (x) accepts a sample x and outputs a scalar value representing the probability that the sample is real data. The goal of the discriminator is to distinguish real samples from generated samples. The discriminator loss function L D consists of two parts: the prediction loss for real samples and the prediction loss for generated samples.

[0028]

[0029] For the "smoky fire" domain, the generator generates samples between "non-smoky fire" and "smoky fire" to help the classifier distinguish fires under different fire intensities, sizes, or background conditions.

[0030] For the "non-smoky fire" domain, the generator generates samples of "slight smoke" or "other similar backgrounds" to enable the model to learn more fine-grained distinctions.

[0031] Step 2, construct a sample set for the UAV fire data model;

[0032] Step 2.1, data annotation;

[0033] Use the labelimg software to annotate the areas with fire in the images and subdivide the fire into two categories: smoke and flame. Through this annotation process, the front and back frame annotation datasets D frame_sample and the time series annotation dataset D seq_sample are generated;

[0034] Step 2.2, adjust the ratio of the number of positive and negative samples;

[0035] Adjust the ratio of the number of positive and negative samples in the front and back frame annotation dataset D frame_sample to ensure the balance of the dataset and avoid biases caused by class imbalance in the model;

[0036] Step 2.3, dataset division;

[0037] Divide the front and back frame annotation dataset D frame_sample into a front and back frame training sample set D frame_train and a front and back frame validation sample set D frame_valWith the front and back frame test sample set D frame_test 。

[0038] To handle the domain differences between "with smoke" and "without smoke", first, the temporal annotation dataset D seq_sample is grouped according to the same GPS coordinates to ensure that historical data at the same location is grouped together. Then, each group of data is divided into two domains according to whether it contains a smoke event: "without smoke" and "with smoke". On this basis, independent feature representations such as image features, background information, texture, and color are introduced for each domain to ensure that the model can learn the feature distributions of these two domains separately. Finally, the dataset is further divided into a temporal training sample set D seq_train , a temporal validation sample set D seq_val and a temporal test sample set D seq_test 。

[0039] Step 3: Construct a single-frame smoke detection model;

[0040] Step 3.1: Train the detection model;

[0041] Install the required software packages and their dependencies, and train the Co-DETR algorithm model on the front and back frame training sample set D frame_train . During training, we used the AdamW optimizer and set a relatively small learning rate of 1e-5;

[0042] The loss function includes classification loss and regression loss. For the classification problem of the predicted bounding boxes, the cross-entropy loss is used to calculate the difference in class labels:

[0043]

[0044] where, is the class probability of the i-th predicted bounding box, and matched is the set of valid predicted bounding boxes after matching through the Hungarian algorithm.

[0045] We use the L1 loss to calculate the regression error of the bounding boxes. The regression error calculates the difference between the center coordinates and dimensions of the bounding box (box) of the predicted bounding box and the ground truth box:

[0046]

[0047] where, t i is the coordinate of the i-th ground truth box, is the coordinate of the i-th predicted bounding box.

[0048] Dynamically adjust the NMS confidence and IOU thresholds according to the size of the target smoke.

[0049] First, calculate the area of the target:

[0050] A = w * h (7) Then, dynamically adjust the IoU threshold:

[0051]

[0052] A threshold is a preset area threshold for distinguishing small targets from large targets. θ min is the IoU threshold for small targets, and θ max is the IoU threshold for large targets.

[0053] Step 3.2, verify the accuracy;

[0054] Perform accuracy verification on the front and back frame verification sample set D frame_val and comprehensively evaluate the performance of the model using evaluation metrics such as accuracy, recall, and F1-score.

[0055] Step 4, construct a temporal fire detection model

[0056] Step 4.1 Image feature extraction and dimensionality reduction

[0057] Extract features from the data in the temporal training sample set D seq_train through multiple convolutional neural networks:

[0058] F l = Conv l (X) = ReLU(Conv(X, K l ) + b l ) (9)

[0059] where K l is the convolutional kernel of the l-th layer, and b l is the bias term, and ReLU is the activation function

[0060] Next, use 1*1 convolution to reduce the dimensionality of the feature map F' to obtain the reduced-dimensional feature map F''

[0061]

[0062] Step 4.2 Vectorization;

[0063] Flatten the multi-dimensional tensor F'' after convolution and dimensionality reduction into a one-dimensional vector for further processing as input to the temporal model:

[0064] v'' seq = {v'' 1 , v'' 2 , …, v'' n} (11)

[0065] where v'' seqRefers to the set of feature vectors of the current data and n-1 historical data.

[0066] Step 4.3 Mutation-aware attention mechanism calculation;

[0067] To handle the mutations (i.e., domain differences) in the "smokeless" and "smoky" time series data, a mutation-aware attention mechanism is introduced. If the data at the current moment is significantly different from the historical data (e.g., changing from smokeless to smoky), a larger weight is assigned to the changing area at the current moment, enabling the model to pay more attention to the changing area at the current moment. First, calculate the correlation between the feature vector Q at the current moment and each feature vector K at the historical moments. This correlation is calculated through the dot product and normalized using scaled dot product attention:

[0068]

[0069] where d k represents the dimension of each feature vector of matrices Q and K.

[0070] Thus, mutation-aware attention and a time-weighted feature map are introduced.

[0071] Step 4.4 Region discrimination;

[0072] Design a separate target region discriminator for the mutation-aware attention and the introduced time-weighted feature map. The target region discriminator consists of a three-layer fully connected neural network. The task of the region discriminator D i,j is to predict the probability that the feature vector at position F i,j is a smoky fire.

[0073] D i,j = σ(W 3 · ReLU(W 2 · ReLU(W 1 · F i,j + b 1 ) + b 2 ) + b 3 ) (13)

[0074] The output of the region discriminator is the probability of whether there is a smoky fire in this region.

[0075] The loss function is the sum of the classification loss and the regression loss:

[0076] L total = 0.6 * L cls + 0.4 * L bbox (14)

[0077] The classification loss uses the cross-entropy loss to measure the gap between the predicted probability and the true label:

[0078]

[0079] The regression loss uses the L1 loss to measure the coordinate difference between the predicted bounding box and the ground truth bounding box:

[0080]

[0081] After combining these two parts of the loss, backpropagation is performed and the model parameters are updated.

[0082] Step 5, joint decision-making;

[0083] Step 5.1, model prediction;

[0084] Input multiple consecutive images in the front and back frame test sample set D frame_test into the model to obtain the firework prediction results and confidence levels for each frame of the image;

[0085] Input the images with the same gps coordinates in the temporal test sample set D seq_test into the model to obtain the firework prediction results and confidence levels for the current temporal image;

[0086] Step 5.2, weighted voting;

[0087] Assign different weights according to the prediction results and confidence levels of the front and back frames and the current temporal image. For the firework prediction result of each frame, use the confidence level of each frame as the weight and combine the results of the front and back frames for decision-making. For the prediction result and confidence level p i of each frame i, sum the prediction results of the front and back frames after weighting:

[0088]

[0089] where k represents the relative position of the front and back frames, p i is the confidence level of the i-th frame, is the prediction result of the i-th frame, and the final prediction result is the weighted voting result.

[0090] Step 6, system deployment;

[0091] Step 6.1, model conversion;

[0092] Convert the trained weight file to the ONNX format, and use C++ and TensorRT to deploy the exported ONNX model on the server for efficient real-time inference.

[0093] Step 6.2, multi-threaded tasks;

[0094] To improve system efficiency, multi-threaded programming is adopted to decouple the video stream reception and inference tasks, preventing the time-consuming inference process from blocking the video stream reception. By creating independent threads to handle the video stream reception and inference tasks respectively, it is ensured that the video stream can be continuously and stably received and saved to the local file system at the frame rate for subsequent inference tasks to use.

[0095] Step 6.3, Dynamically transmit the video stream;

[0096] During the process of pushing the video stream, the network status is monitored in real time. Through FFMPEG, the bandwidth change is monitored and the bit rate and resolution of the video stream are adjusted. The adaptive bit rate streaming (ABR) function is adopted to select appropriate video coding settings (such as H.264, H.265, etc.) according to the real-time network status. A low-resolution video stream is provided when the bandwidth is low, and a high-definition video stream is provided when the bandwidth is high. In addition, by monitoring the transmission delay and packet loss rate in real time, the resolution is reduced when the network quality is poor to ensure transmission stability.

[0097] Step 7, Model optimization;

[0098] Step 7.1, Collect data;

[0099] During the operation of the system, for the false alarm and missed alarm data generated, they are collected, analyzed, and labeled to form a supplementary dataset D add .

[0100] Step 7.2, Incremental learning;

[0101] Based on the existing model parameters, the supplementary dataset D add is introduced to retrain the model to improve its performance. Here, the incremental learning method of gradient descent is adopted:

[0102]

[0103] θ t is the current model parameter, θ t+1 is the model parameter updated through incremental learning, η is the learning rate used to control the step size of parameter update, is the gradient of the loss function with respect to the parameter, indicating the sensitivity of the error of the model under the current parameter θ t to the supplementary dataset D add .

[0104] The second aspect of the present invention relates to an automatic inspection device based on drones, including a plurality of drone nests, a plurality of drones, and a memory. The memory stores executable code. When the one or more processors execute the executable code, it is used to implement the method for fire recognition and positioning in the drone inspection of the present invention.

[0105] The third aspect of the present invention relates to a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the method for fire recognition and positioning in the drone patrol inspection of the present invention.

[0106] The present invention aims to achieve real-time monitoring, precise detection, and rapid response to the fire phenomena within the patrol inspection area. First, a fusion dataset of true-color and infrared images of drone fires and a time-series dataset under the same coordinates are established; then, a generative adversarial model is used for data augmentation; next, a CO-DETR model is used to construct a fusion image model; at the same time, a time-series model is designed, covering image feature extraction and dimensionality reduction, vectorization processing of time-series data, attention calculation for mutation perception, and region discrimination; a weighted voting mechanism is adopted to weight the prediction results according to the confidence levels of each model; the video stream reception and inference tasks are decoupled through multi-threading technology, while listening to the network status and dynamically adjusting the video stream transmission; finally, an incremental learning method is used to iteratively optimize the model.

[0107] The advantages of the present invention are as follows:

[0108] 1. The present invention can achieve rapid response, ensuring accurate recognition and timely handling of the fire phenomena within the patrol inspection area.

[0109] 2. The fusion dataset combines different spectral information, improving the accuracy in complex environments, especially in low-light or high-temperature situations.

[0110] 3. Using a generative adversarial network (GAN) for data augmentation solves the problem of insufficient training data. By generating more diverse training samples, the generalization ability of the model is enhanced.

[0111] 4. Adopting a weighted voting mechanism to weight the prediction results according to the confidence levels of each model can improve the accuracy and stability of the prediction results, avoiding the risk of misjudgment by a single model.

[0112] 5. By decoupling the video stream reception and inference tasks through multi-threading technology, large-scale video data streams can be efficiently processed. At the same time, listening to the network status and dynamically adjusting the video stream transmission effectively improves the overall performance and response speed of the system.

[0113] 6. Using an incremental learning method enables the model to be iteratively optimized while continuously receiving new data. This method allows the model to gradually improve its accuracy and robustness during long-term operation, adapting to new environmental changes and fire detection tasks. Description of the Drawings

[0114] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings involved in the implementation process. It should be understood that the following drawings only show some embodiments of the present invention and do not limit the scope of the invention. For ordinary technicians in this field, other related drawings can be derived based on these drawings without creative work.

[0115] Figure 1 The present invention is a flowchart of a smoke and fire detection system for implementing the method of the present invention.

[0116] Figure 2 It is a flow chart of the dynamic resolution transmission technology of the present invention.

[0117] Figure 3 Schematic diagram of the target area discriminator of the present invention.

[0118] Figure 4 It is a schematic diagram of the smoke and fire detection and recognition effect of the present invention.

[0119] Figure 5 It is a schematic diagram of the device of the present invention. DETAILED DESCRIPTION

[0120] The following will be combined with the drawings in the embodiments of the present application to describe the technical solution in detail and clearly. It should be noted that the described embodiments are only part of the present application and do not represent all implementation methods. The various components and their configurations shown in the drawings can be arranged and designed differently according to actual needs, with certain flexibility and variability.

[0121] Example 1

[0122] Reference Figure 1 , Figure 2 ,The method for identifying and locating fireworks in drone inspection includes the following steps:

[0123] Step 1: Perform inspection tasks;

[0124] Step 1.1, publish the task;

[0125] The drone unified management platform receives and processes task information according to the requirements of the inspection task, and assigns the task to the corresponding drone terminal. The task content includes information such as the designated inspection area, monitoring target and task time, ensuring that the drone can efficiently perform the inspection work according to the predetermined plan.

[0126] Step 1.2, inspect the target area;

[0127] After receiving the task, the drone starts and flies along the set route, entering the designated inspection area. After reaching the inspection area, the drone begins to capture real-time videos of the area and perform smoke and fire detection tasks. The drone transmits the captured video stream to the background system and simultaneously sends a smoke and fire detection request to the system.

[0128] Step 2, Smoke and Fire Analysis;

[0129] Step 2.1, Model Inference;

[0130] When the system receives the smoke and fire detection request, it starts to receive the real-time video stream transmitted by the drone. The system then automatically activates the trained smoke and fire detection algorithm Co-DETR to perform inference analysis on the real-time captured images and identify potential flame or smoke areas in the images.

[0131] Meanwhile, the system retrieves historical data at the same location to construct temporal input features. Through joint inference of the historical data and the current image by the temporal model, the system can determine whether there is a smoke and fire area in the current image.

[0132] Step 2.2, System Voting;

[0133] After receiving the detection results of multiple consecutive frames and temporal sequences, the system uses the prediction results and confidence levels to perform joint decision-making using the weighted voting method. By comprehensively analyzing the detection results of multiple frames, the system can more accurately determine whether there is a smoke and fire phenomenon. This method effectively reduces the risk of false judgment in a single frame, thus significantly reducing the occurrence of false alarms and improving the accuracy and reliability of smoke and fire detection.

[0134] Step 3, Result Transmission;

[0135] Step 3.1, Manual Review;

[0136] When the system detects a smoke and fire area, it automatically marks the area as an alarm event and transmits the color image of the smoke and fire to the management platform in real time via the network, providing the staff with the real-time situation on site. After receiving the warning information, the management platform immediately initiates the manual review process. The staff further analyzes the transmitted images and detection data to confirm the authenticity of the fire. If a fire or fire situation is confirmed, the platform will promptly notify the relevant emergency departments and coordinate the actual fire-fighting operations to ensure timely response and handling.

[0137] Step 3.2, Video Push;

[0138] The images processed by the system will be transmitted to the management platform for display in the form of a video stream. During the pushing process, the system will monitor the network status in real time and automatically adjust the resolution of the video or image according to the current network bandwidth and stability to ensure the stability and efficiency of the transmission process. If the network condition is good, the system will transmit high-resolution images or videos to retain more details; while in the case of limited bandwidth or unstable network, the system will automatically reduce the resolution to reduce the amount of data transmission, thus avoiding the interruption of the video stream caused by network latency or packet loss and ensuring the timely transmission of key data.

[0139] Step 4, cyclic reasoning;

[0140] In the case of completing the sending of early warning information or not detecting smoke and fire currently, the system will continuously perform reasoning and analysis on the video stream until the end of the video stream. Through this cyclic reasoning mechanism, the system can achieve full-process dynamic monitoring of the monitored area to ensure the comprehensiveness and real-time nature of the patrol task.

[0141] Step 5, model optimization strategy;

[0142] Step 5.1, collection of false alarm and missed alarm data;

[0143] For the cases of false alarms and missed alarms confirmed by manual review, the system will automatically collect this data and classify, organize, and label it regularly. By integrating these labeled data, a new supplementary sample set is formed to provide data support for subsequent model optimization.

[0144] Step 5.2, continuous optimization and reinforcement learning;

[0145] The system adopts a reinforcement learning strategy to continuously optimize the flame detection model. Combining the new data and improved algorithms, the model parameters are continuously updated through incremental learning to gradually reduce the phenomena of false alarms and missed alarms. Through this continuous optimization process, the detection accuracy and robustness of the system are significantly improved, providing a more reliable fire monitoring ability for practical applications.

[0146] Embodiment 2

[0147] Refer to Figure 5 , this embodiment relates to an automatic inspection device based on an unmanned aerial vehicle (UAV), including a plurality of UAV nests, a plurality of UAVs, and a memory. An executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the method for smoke and fire recognition and positioning in the UAV inspection of Embodiment 1.

[0148] Embodiment 3

[0149] This embodiment relates to a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the method for smoke and fire recognition and positioning in the UAV inspection of Embodiment 1.

[0150] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the device embodiments, the above description is only the preferred embodiment of the present invention. Since it is basically similar to the method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the related content. As described above, the above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. For any person skilled in the art in the technical field disclosed by the present invention, for those of ordinary skill in the technical field, any changes or substitutions that can be easily thought of without departing from the principle of the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. The method for identifying and locating fireworks in drone inspections specifically includes the following steps: Step 1: Prepare data, including making previous and next frame data sets and current position history data sets and performing data enhancement; Step 2, create a sample set of drone fireworks data model; Step 3, construct a single-frame fireworks detection model; Step 4, constructing a time series fireworks detection model; Step 5: Execute joint decision; Step 6: System deployment and integration; Step 7: Model optimization and iterative update.

2. The method for identifying and locating fireworks in drone inspection according to claim 1, characterized in that: Step 1 includes: Step 1, create a data set; Step 1.1, front and back frame data set; The drone inspection video is processed by frame extraction, and the infrared image obtained by the infrared sensor is used as the fourth channel and fused into the extracted image frame to generate the inspection image dataset; the current frame is used as a reference to obtain the adjacent front and back frames to form the front and back frame dataset D frame : D frame ={P1,P2,…,P n-k ,…,P n } (1) Where P i , i = 1, 2, ..., nk, ..., n represents the image data of the nkth frame and its preceding and following frames in the video; each frame of the image contains RGB channels and an infrared image as the fourth channel; Step 1.2, time series data set; During the flight of the drone, the system records the GPS coordinates and timestamp of the current location; associates the historical data of the same location to create a time series data set D seq : D seq ={(P1,(lon1,lat1,alt1),t1),(P2,(lon2,lat2,alt1),t2),…, (P n ,(lon n ,lat n ,alt n ),t n )} (2) Among them, P i is the i-th frame image data, (lon i ,lat i ,alt i ) is the coordinate when the i-th frame image was taken, t i is the timestamp of the i-th frame image; Step 1.3, enhance data; Perform data enhancement on the collected images, including conventional methods such as image flipping, translation, scaling, and adding noise. Mosaic data enhancement technology is also introduced to generate more training samples. In order to further improve the model's adaptability to the differences between "no fireworks" and "with fireworks" domains, a generative adversarial network is used to generate data samples with different domain characteristics, thereby enhancing the model's ability to adapt to differences between domains; the generative adversarial network includes a generator and a discriminator; The generator G samples a vector z~p from the latent space z , where pz is a Gaussian distribution in the latent space; the generator then uses this vector z as input to generate a sample G θ (z); by generating samples G θ (z) Let the discriminator D φ Considering that these samples are as real as possible, the output D of the discriminator for the generated samples φ (G θ (z)) is close to 1; the generator loss function L G Defined as: Discriminator D φ (x) accepts a sample x and outputs a scalar value indicating the probability that the sample is real data; the goal of the discriminator is to distinguish between real samples and generated samples; the loss function of the discriminator is L D It consists of two parts: the prediction loss of real samples and the prediction loss of generated samples; For the "fireworks" domain, the generator generates samples between "no fireworks" and "fireworks", helping the classifier to distinguish fireworks under different firework intensities, sizes or background conditions; For the "no smoke and fire" domain, the generator generates samples of "light smoke" or "other similar backgrounds" to allow the model to learn more fine-grained distinctions.

3. The method for identifying and locating fireworks in drone inspection according to claim 1, characterized in that: Step 2 includes: Step 2.1, data annotation; Use labelimg software to mark the area where fireworks appear in the image, and subdivide the fireworks into two categories: smoke and flame. Through this marking process, the front and back frame annotation dataset D is generated. frame_sample And the time series annotation dataset D seq_sample ; Step 2.2, adjust the ratio of positive and negative samples; Label the dataset D for the previous and next frames based on the ratio of positive and negative samples frame_sample Adjust to ensure the balance of the data set and avoid bias caused by imbalanced categories in the model; Step 2.3, data set division; The previous and next frames are labeled as dataset D frame_sample Divide into the previous and next frame training sample set D according to a certain ratio frame_train , front and back frame verification sample set D frame_val And the previous and next frame test sample set D frame_test ; In order to deal with the domain differences between "with fireworks" and "without fireworks", we first label the time series dataset D seq_sample The data are grouped according to the same GPS coordinates to ensure that the historical data at the same location are grouped together. Then, each group of data is divided into two areas according to whether it contains fireworks events: "no fireworks" and "with fireworks". On this basis, independent feature representations such as image features, background information, texture and color are introduced for each area to ensure that the model can learn the feature distribution of these two areas separately. Finally, the data set is further divided into a time series training sample set D according to a certain ratio. seq_train , Timing Verification Sample Set D seq_val With the timing test sample set D seq_test .

4. The method for identifying and locating fireworks in drone inspection according to claim 1, characterized in that: Step 3 includes: Step 3.1, training the detection model; Install the required packages and their dependencies, and train the sample set D on the previous and next frames. frame_train The Co-DETR algorithm model was trained using the AdamW optimizer and a smaller learning rate of 1e-5 was set. The loss functions include classification loss and regression loss. For the classification problem of the prediction box, the cross entropy loss is used to calculate the difference in category labels: in, is the category probability of the i-th prediction box, and matched is the set of valid prediction boxes after matching by the Hungarian algorithm; Use L1 loss to calculate the regression error of the box; the regression error calculates the difference between the coordinates and size of the center point of the bounding box box of the predicted box and the true box: Among them, t i is the coordinate of the i-th ground-truth box, is the coordinate of the i-th prediction box; Dynamically adjust the NMS confidence and IOU threshold according to the target fireworks size; First, calculate the area of ​​the target: A=w*g (7) Then, dynamically adjust the IoU threshold: A threshold is a preset area threshold used to distinguish small targets from large targets; θ min is the IoU threshold of small objects, θ max is the IoU threshold of large objects; Step 3.2, verify accuracy; Verify the sample set D in the previous and next frames frame_val The accuracy is verified on the dataset, and the performance of the model is comprehensively evaluated using evaluation indicators such as precision, recall rate, and F1 score.

5. The method for identifying and locating fireworks in drone inspection according to claim 1, characterized in that: Step 4: The specific steps for building a time series fireworks detection model are as follows: Step 4.1 Image feature extraction and dimensionality reduction; The time series training sample set D seq_train The data in is passed through multiple convolutional neural networks to extract features: F l =Conv l (X)=ReLU(Conv(X,K l )+b l ) (9) Among them, K l is the convolution kernel of layer l, b l is the bias term, ReLU is the activation function. Next, use 1*1 convolution to reduce the dimension of the feature map F′ to obtain the reduced dimension feature map F″ Step 4.2 vectorization; The multi-dimensional tensor F″ after convolution and dimensionality reduction is flattened into a one-dimensional vector so that it can be provided as input to the time series model for further processing: v″ seq ={v″1,v″2,…,v″ n } (11) where v″ seq Refers to the feature vector set of current data and n-1 historical data; Step 4.3 mutation-aware attention mechanism calculation; In order to deal with the mutations, i.e., domain differences, in the "no fireworks" and "with fireworks" time series data, a mutation-aware attention mechanism is introduced; if the data at the current moment is significantly different from the historical data, a larger weight is given to the change area at the current moment, so that the model pays more attention to the change area at the current moment; first, the correlation is calculated between the feature vector Q at the current moment and each feature vector K at the historical moment; the correlation is calculated by dot product and normalized using scaled dot product attention: Where d k Represents the dimension of each eigenvector of matrices Q and K; This results in mutation-aware attention and the introduction of time-weighted feature maps; Step 4.4: Region identification; A target region discriminator is designed for mutation-aware attention and the introduction of time-weighted feature maps. The target region discriminator consists of a three-layer fully connected neural network. i,j The task is to predict the position F i,j The probability that the feature vector at is a firework; D i,j =σ(W3·ReLU(W2·ReLU(W1·F i,j +b1)+b2)+b3) (13) The output of the region discriminator is whether the region contains fireworks and the probability of fireworks; The loss function uses the sum of classification loss and regression loss: L total =0.6*L cls +0.4*L bbox (14) The classification loss uses the cross entropy loss to measure the difference between the predicted probability and the actual label: The regression loss uses L1 loss to measure the coordinate difference between the predicted box and the real box: After combining these two parts of loss, backpropagation is performed and the model parameters are updated.

6. The method for identifying and locating fireworks in drone inspection according to claim 1, characterized in that: Step 5 includes: Step 5, joint decision making; Step 5.1, model prediction; The previous and next frame test sample set D frame_test Multiple frames of continuous images are input into the model to obtain the fireworks prediction results and confidence of each frame of image; The timing test sample set D seq_test The same GPS coordinate image is input into the model to obtain the fireworks prediction result and confidence of the current time series image; Step 5.2, weighted voting; Different weights are assigned according to the prediction results and confidence levels of the previous and next frames and the current time series image; for the fireworks prediction results of each frame, the confidence level of each frame is used as a weight, and the decision is made in combination with the results of the previous and next frames; for the prediction results and confidence level p of each frame i, i , weighted sum of the prediction results of the previous and next frames: Among them, k represents the relative position of the previous and next frames, p i is the confidence of the i-th frame, is the prediction result of the i-th frame, and the final prediction result is the weighted voting result.

7. The method for identifying and locating fireworks in drone inspection according to claim 1, characterized in that: Step 6 includes: Step 6.1, model conversion; Convert the trained weight file to ONNX format, and use C++ and TensorRT to deploy the exported ONNX model on the server for efficient real-time reasoning; Step 6.2, multi-threaded tasks; To improve system efficiency, multi-threaded programming is used to decouple video stream reception from inference tasks, preventing the time-consuming inference process from blocking video stream reception. Independent threads are created to handle video stream reception and inference tasks respectively, ensuring that the video stream can be continuously and stably received and saved to the local file system at the frame rate for use in subsequent inference tasks. Step 6.3, dynamically transmit video stream; During the process of pushing video streams, the network status is monitored in real time, bandwidth changes are monitored through FFMPEG and the bit rate and resolution of the video stream are adjusted. The adaptive bit rate stream ABR function is used to select appropriate video encoding settings according to the real-time network status, providing low-resolution video streams when the bandwidth is low and high-definition video streams when the bandwidth is high. In addition, by monitoring the transmission delay and packet loss rate in real time, the resolution is reduced when the network quality is poor to ensure transmission stability.

8. The method for identifying and locating fireworks in drone inspection according to claim 1, characterized in that: Step 7: The specific steps of model optimization and iterative update are as follows: Step 7, model optimization; Step 7.1, collect data; During the operation of the system, the false positive and false negative data generated are collected, analyzed and annotated to form a supplementary data set D add ; Step 7.2, incremental learning; Based on the existing model parameters, a supplementary dataset D is introduced. add , retrain the model to improve performance; here we adopt the incremental learning method of gradient descent: θ t is the current model parameter, θ t+1 is the model parameter updated by incremental learning, η is the learning rate used to control the step size of parameter update, is the gradient of the loss function with respect to the parameter, indicating that the model is at the current parameter θ t The error under the supplementary data set D add sensitivity.

9. An automatic inspection device in drone inspection, comprising multiple drone nests, multiple drones, and a memory, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used to implement the fireworks recognition and positioning method in drone inspection of any one of claims 1-7.

10. A computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the method for identifying and locating fireworks in drone inspections according to any one of claims 1 to 7.

Citation Information

Cited By

  • Industrial robot motion control method, system, equipment and medium

    CN120347776A

  • Smoke and fire point identification method

    CN120876979A

  • Method and device for discriminating behavior event of aircraft

    CN121477150A