Auxiliary fire extinguishing method and system based on unmanned aerial vehicle information and image processing technology

By improving the YOLOv8 model and image smoke removal technology, combined with the drone gimbal position information, the problems of drone image target detection and three-dimensional positioning in fire protection scenarios are solved, and efficient fire resource scheduling and intelligent rescue are achieved.

CN120355899AActive Publication Date: 2025-07-22CHINA UNIV OF MINING & TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510467843.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-22
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

In fire protection scenarios, the image of the drone is affected by flame, smoke and water jets, and the traditional atmospheric scattering model is not applicable. Drones equipped with monocular cameras lack effective three-dimensional visual positioning methods, which leads to difficulty in target detection and positioning, affecting fire rescue efficiency.

Method used

The improved YOLOv8 model is used for object detection, and combined with UNet and ResNet networks for image desmog treatment, the Berlin noise generation model is used to generate smoke images, and the improved atmospheric scattering model is pre-processed, and visual target positioning is combined with the drone gimbal positioning information to achieve three-dimensional positioning.

Benefits of technology

It provides richer fire scene information, assists in scheduling fire resources, improves the intelligence level and efficiency of fire rescue, and reduces costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355899A_ABST
    Figure CN120355899A_ABST
Patent Text Reader

Abstract

The invention discloses an auxiliary fire extinguishing method and system based on unmanned aerial vehicle information and an image processing technology. The method comprises the steps that S1, an unmanned aerial vehicle carrying preset equipment is used for collecting a fire scene image; s2, performing smoke removal preprocessing on the fire scene image to obtain a preprocessed image; s3, performing target detection on the preprocessed image to obtain a target detection result; s4, performing visual target positioning of the fire scene based on the holder pose information of the unmanned aerial vehicle and the target detection result; and S5, based on the target detection result and the visual target positioning, auxiliary scheduling of fire-fighting resources is carried out, and auxiliary fire extinguishing based on unmanned aerial vehicle information and an image processing technology is completed. And carrying out target detection on the preprocessed image by using an improved YOLOv8 model, and carrying out three-dimensional positioning on a fire source target based on a camera imaging principle in combination with a target detection result and unmanned aerial vehicle monocular holder sensor data so as to assist in dispatching fire-fighting resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of fire fighting and rescue and image processing, and particularly relates to an auxiliary fire extinguishing method and system based on unmanned aerial vehicle information and image processing technology. Background Art

[0002] The occurrence of fire accidents often causes serious casualties and property losses, has a huge impact on society and economy, and also poses a great threat to firefighters. Improving the intelligent level of fire fighting and rescue work, increasing the efficiency and safety factor of fire fighting and rescue work is an urgent problem to be solved. In fire fighting and rescue work, it is crucial to obtain fire scene information safely and quickly. Unmanned aerial vehicles (UAVs) have the advantages of being flexible, low-cost, rich in functions, and strong in scalability, and have great application value in fire fighting scenarios. UAVs can carry various sensors and devices, such as data transmission modules and image transmission modules, quickly reach the fire scene, communicate with the fire fighting command terminal, and provide real-time fire scene images for fire fighting and rescue personnel. Using image processing technology and intelligent algorithms to integrate UAV images and sensor information can obtain information such as the location of the fire source and the size of the fire, so as to assist firefighters in dispatching fire fighting resources and carrying out efficient and safe fire extinguishing and disaster relief.

[0003] In fire fighting scenarios, affected by flames, smoke, and water jets, the UAV flies at a higher altitude and has a wider field of view, resulting in a higher proportion of small targets in the captured images. At the same time, the images contain non-uniform light and smoke generated by the flames. In the traditional atmospheric scattering model, the pixel value of aerosol is defaulted to (255, 255, 255), which is only applicable to scenes with uniform light intensity, so it is not applicable to fire fighting scenarios. In view of this, it is urgent to study a target detection scheme for UAV images applicable to fire fighting scenarios.

[0004] Most traditional three-dimensional vision positioning methods rely on binocular vision systems, use feature matching algorithms to generate disparity maps or point clouds, and then perform three-dimensional positioning on the targets in the images. Compared with binocular vision systems, UAVs equipped with monocular vision systems are more portable and lower in cost. Therefore, in this design, a UAV equipped with a monocular camera is adopted. At this time, the traditional binocular vision positioning method fails, and it is urgent to study a three-dimensional vision positioning method for UAVs equipped with monocular cameras. Summary of the Invention

[0005] In order to solve the above problems, the present invention proposes an auxiliary fire extinguishing method based on UAV information and image processing technology. The method specifically includes:

[0006] Step S1: Use a UAV equipped with preset equipment to collect fire scene images;

[0007] Step S2: Perform de-smoking preprocessing on the fire scene images to obtain preprocessed images;

[0008] Step S3: Perform object detection on the preprocessed image to obtain an object detection result;

[0009] Step S4: Based on the pan-tilt pose information of the drone and the object detection result, perform visual target localization at the fire scene;

[0010] Step S5: Based on the object detection result and the visual target localization, assist in dispatching fire resources to complete fire extinguishing assistance based on drone information and image processing technology.

[0011] Optionally, in step S2, the process of obtaining the preprocessed image specifically includes:

[0012] Construct a neural network model based on the UNet network and the ResNet network. During neural network training, use the Berlin noise to generate model parameters t(x) and c(x), input the model parameters t(x) and c(x) into the improved atmospheric scattering model to generate a smoke image, and use the smoke image and the existing fire scene images to train the neural network model to obtain a trained neural network model;

[0013] Input the fire scene image into the trained neural network model to obtain the preprocessed image.

[0014] Optionally, the process of constructing the neural network model based on the UNet network and the ResNet network specifically includes:

[0015] The neural network model includes an encoder network and a decoder network;

[0016] The encoder network uses the ResNet network to extract features. The ResNet network contains 5 neural network layers. The first layer consists of a convolutional layer, a batch normalization layer, a ReLU activation function layer, and a max pooling layer in sequence. The input image is resized to a dimension size of 640×640×3, and a feature map of 320×320×3 is output. The second, third, and fourth layers are residual network layers, and the residual network layer consists of a convolutional layer and a residual connection;

[0017] The decoder network consists of 5 neural network layers. The first layer receives the output of the encoder. The first, second, third, and fourth layers consist of a transposed convolutional layer, a channel concatenation layer, and a residual network layer. The input of each layer first passes through a transposed convolution, and then passes through the channel concatenation layer to perform channel concatenation with the feature map of the same dimension size in the encoder network, and finally passes through the residual network layer to extract features and output. The fifth layer consists of two convolutional layers, receives the feature map of 640×640×32 output by the fourth layer of the decoder network, and outputs a dimension size of 640×640×3 to obtain the decoder output result.

[0018] Optionally, the process of using the Berlin noise to generate the model parameters t(x) and c(x) and inputting the model parameters t(x) and c(x) into the improved atmospheric scattering model to generate a smoke image specifically includes:

[0019] The generation process of the t(x) specifically includes:

[0020] Set the number of noise layers to 6, the persistence to 0.5, the porosity to 2.0, the scale to 250, and generate a Berlin noise map P after normalization noise ;

[0021] Based on the Berlin noise map, generate a smoke transmittance map t(x):

[0022] t(x) = P noise ·(t max - t min ) + t min

[0023] where t max is the maximum value of the transmittance, and t min is the minimum value of the transmittance;

[0024] The generation process of the c(x) specifically includes:

[0025] Based on the Berlin noise map, generate the RGB value c(x) of the smoke haze:

[0026] c(x) = P noise ·(c max - c min ) + c min ;

[0027] where c max is the maximum value of the RGB of the smoke haze, and c max is the minimum value of the RGB of the smoke haze;

[0028] Substitute the obtained t(x) and c(x) into the improved atmospheric scattering model to add smoke to the clear image and obtain a smoke image.

[0029] Optionally, the improved atmospheric scattering model is specifically:

[0030] I(x) = J(x)t(x) + c(x)(1 - t(x))

[0031] where I(x) is the fire scene image; J(x) is the image after ideal defogging processing; t(x) is the smoke transmittance map; c(x) is the RGB value of the smoke haze.

[0032] Optionally, in step S4, the process of visually locating the target at the fire scene specifically includes:

[0033] According to the camera calibration method, obtain the camera internal parameter matrix, and perform distortion processing on the camera based on the camera internal parameter matrix;

[0034] Use the orthodontic UAV positioning module to obtain the position information of the UAV;

[0035] Obtain the attitude information of the UAV gimbal camera based on the orthodontic UAV sensor;

[0036] Obtain the translation vector based on the position information of the UAV, calculate the rotation matrix based on the attitude information of the UAV camera gimbal, and obtain the imaging external parameter matrix;

[0037] Based on the target detection result, the imaging external parameter matrix, and the camera imaging principle, realize the conversion of the pixel coordinate system - camera coordinate system - world coordinate system, and obtain the initial coordinate solution of the target to be located in the imaging plane under the world coordinate;

[0038] Use bundle adjustment to optimize the three-dimensional coordinates of the target to be located according to the observation results of each pose of the UAV, and obtain the visual target location at the fire scene.

[0039] The present invention also discloses an auxiliary fire extinguishing system based on UAV information and image processing technology. The system includes:

[0040] An image acquisition module, which is used to acquire fire scene images using a UAV equipped with preset equipment;

[0041] An image preprocessing module, which is used to perform de-smoke preprocessing on the fire scene images to obtain preprocessed images;

[0042] A target detection module, which is used to perform target detection on the preprocessed images to obtain target detection results;

[0043] A target location module, which is used to perform visual target location at the fire scene based on the gimbal pose information of the UAV and the target detection result;

[0044] A fire extinguishing assistance module, which is used to assist in scheduling fire fighting resources based on the target detection result and the visual target location, and complete the auxiliary fire extinguishing based on UAV information and image processing technology.

[0045] Optionally, the process of obtaining the preprocessed image specifically includes:

[0046] Build a neural network model based on the UNet network and the ResNet network. During the training of the neural network, use the Berlin noise to generate the model parameters t(x) and c(x), input the model parameters t(x) and c(x) into the improved atmospheric scattering model to generate a smoke image, and use the smoke image and the existing fire scene image to train the neural network model to obtain a trained neural network model;

[0047] Input the fire scene image into the trained neural network model to obtain a preprocessed image.

[0048] Optionally, the process of building the neural network model based on the UNet network and the ResNet network specifically includes:

[0049] The neural network model includes an encoder network and a decoder network;

[0050] The encoder network uses the ResNet network to extract features. The ResNet network contains 5 neural network layers. The first layer consists of a convolutional layer, a batch normalization layer, a ReLU activation function layer, and a max pooling layer in sequence. The input image is resized to a dimension of 640×640×3, and a feature map of 320×320×3 is output. The second, third, and fourth layers are residual network layers, and the residual network layer consists of a convolutional layer and a residual connection;

[0051] The decoder network consists of 5 neural network layers. The first layer receives the output of the encoder. The first, second, third, and fourth layers consist of a transposed convolutional layer, a channel concatenation layer, and a residual network layer. The input of each layer first passes through a transposed convolution, and then passes through the channel concatenation layer to be concatenated with the feature map of the same dimension size in the encoder network, and finally passes through the residual network layer to extract features and output. The fifth layer consists of two convolutional layers, receives the feature map of 640×640×32 output by the fourth layer of the decoder network, and the output dimension size is 640×640×3 to obtain the decoder output result.

[0052] Optionally, the process of using the Berlin noise to generate the model parameters t(x) and c(x), and inputting the model parameters t(x) and c(x) into the improved atmospheric scattering model to generate a smoke image specifically includes:

[0053] The generation process of t(x) specifically includes:

[0054] Set the number of noise layers to 6, the persistence to 0.5, the porosity to 2.0, the ratio to 250, and generate the Berlin noise map P after normalization noise ;

[0055] Based on the Berlin noise map, generate the smoke transmittance map t(x):

[0056] t(x) = P noise ·(t max - t min ) + t min

[0057] where t max is the maximum value of the transmittance, and t min is the minimum value of the transmittance;

[0058] The generation process of the said c(x) specifically includes:

[0059] Based on the said Perlin noise map, generate the RGB values c(x) of the smoke and haze:

[0060] c(x) = P noise ·(c max - c min ) + c min ;

[0061] where c max is the maximum value of the RGB of the smoke and haze, and c max is the minimum value of the RGB of the smoke and haze;

[0062] Substitute the t(x) and c(x) obtained above into the improved atmospheric scattering model to add smoke to the clear image to obtain a smoke image.

[0063] Optionally, the improved atmospheric scattering model is specifically:

[0064] I(x) = J(x)t(x) + c(x)(1 - t(x))

[0065] where I(x) is the fire scene image; J(x) is the image after ideal defogging processing; t(x) is the smoke transmittance map; c(x) is the RGB value of the smoke and haze.

[0066] Compared with the prior art, the beneficial effects of the present invention are:

[0067] The present invention takes into account the characteristics of the images obtained by drones in the fire scene, proposes an image pre - processing method for removing smoke, and uses the improved YOLOv8 model to perform object detection on the pre - processed images. Based on the camera imaging principle, combined with the object detection results and the data of the drone's monocular gimbal sensor, three - dimensional positioning of the fire source target is carried out to assist in dispatching fire - fighting resources. The present invention provides richer fire - scene information for the fire - fighting and rescue process at a lower cost, assists in dispatching fire - fighting resources, and improves the intelligent level of fire - fighting work. Brief Description of the Drawings

[0068] To more clearly illustrate the technical solution of the present invention, the following briefly introduces the attached drawings required in the embodiments. Obviously, the attached drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other attached drawings can be obtained based on these attached drawings.

[0069] Figure 1 FIG. 4 is a flowchart of the method for the auxiliary fire extinguishing method based on unmanned aerial vehicle information and image processing technology according to an embodiment of the present invention;

[0070] Figure 2 FIG. 5 is a schematic diagram of the network model of the image dehazing algorithm proposed in an embodiment of the present invention;

[0071] Figure 3 FIG. 6 is a schematic diagram of the improved YOLOv8 network model used in an embodiment of the present invention. Detailed Embodiments

[0072] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the attached drawings and specific embodiments.

[0073] Embodiment 1

[0074] An auxiliary fire extinguishing method based on unmanned aerial vehicle information and image processing technology, as Figure 1 shown, the method includes:

[0075] Step S1: Use a drone equipped with preset equipment to collect fire scene images.

[0076] Before the drone takes off, it needs to be equipped with a calibrated pan-tilt camera, various sensors, and a wireless communication module, so that the ground terminal device can obtain the three-dimensional position, attitude information of the drone pan-tilt camera, and the drone video stream in real time. The drone flies over the fire scene to monitor the fire scene.

[0077] Step S2: Perform de-smoke preprocessing on the fire scene images to obtain preprocessed images.

[0078] As Figure 2 shown, according to the classic atmospheric scattering model in the field of image dehazing, aiming at suppressing the smoke in the fire scene images, the improved formula is as follows:

[0079] I(x) = J(x)t(x) + c(x)(1 - t(x))

[0080] Wherein, I(x) is the image before preprocessing; J(x) is the image after ideal haze removal processing; t(x) is the smoke transmittance map, which can reflect the relative smoke concentration, and the smaller the value of t(x) in the area with higher smoke concentration; c(x) is the RGB value of the smoke and haze. t(x) and c(x) are obtained through an encoder neural network and two decoder neural networks. The input image to be processed first passes through the encoder neural network to obtain a feature map with a dimension of 40×40×512, which is used as the input of the decoder network. The two decoders respectively output t(x) and c(x). Substituting the decoder output into the improved haze imaging model formula, we get

[0081]

[0082] The encoder and decoder neural networks use UNet and ResNet in combination. The encoder network extracts features based on ResNet and consists of 5 neural network layers. The first layer is sequentially composed of a convolutional layer, a batch normalization layer, a ReLU activation function layer, and a max pooling layer. The input image is Resized to a dimension of 640×640×3, and a feature map of 320×320×3 is output. The following 4 layers are residual network layers, which are composed of convolutional layers and residual connections. The 4 residual network layers sequentially output feature maps with dimensions of 320*320*64, 160*160*128, 80*80*256, and 40*40*512; the decoder network consists of 5 neural network layers. The first layer receives the output of the encoder. The structures of the first four layers are the same and are composed of a transposed convolutional layer, a channel concatenation layer, and a residual network layer. The input of each layer first passes through the transposed convolutional layer, and then passes through the channel concatenation layer to concatenate the channels with the feature map of the same dimension in the encoder network. Finally, features are extracted and output through the residual network layer. The fifth layer of the decoder network is composed of two convolutional layers, receives the feature map of 640×640×32 output by the fourth layer, and outputs a dimension of 640×640×3. The two decoders respectively output t(x) and c(x).

[0083] During the training process of the neural network, some of the smoke images used are generated based on Perlin noise and an improved atmospheric scattering model. Specifically, it includes the following steps:

[0084] (1) Generate a Perlin noise map. The number of noise layers octaves is set to 6, the Persistence is set to 0.5, the lacunarity is set to 2.0, and the scale is set to 250. Normalize the generated Perlin noise map to obtain P noise .

[0085] (2) Use Perlin noise to generate the transmittance map t(x). Let the maximum value t max of the transmittance map t(x) be 0.75, and the minimum value tmin is 0.1, and the transmittance map t(x) is generated by the formula t(x) = P noise ·(t max -t min ) + t min

[0086] (3) Generate the smoke RGB map c(x) using Perlin noise. Let the maximum value c max of c(x) be 255 and the minimum value c min be 120. The c(x) is generated by the formula c(x) = P noise ·(c max -c min ) + c min

[0087] (4) Substitute the obtained t(x) and c(x) into the improved atmospheric scattering model to obtain the image with increased haze.

[0088] Step S3: Perform object detection on the preprocessed image to obtain the object detection result.

[0089] As Figure 3 shown, for object detection, an improved YOLOv8 architecture is adopted. A 160×160 object detection head is added to the YOLOv8 architecture to enhance the object detection ability; the ordinary upsampling is changed to dynamic upsampling, which is more flexible than the upsampling methods of nearest neighbor and bilinear interpolation; the attention mechanism is introduced into the YOLOv8 network structure, and FasterNetBlock is added to the shallow layer of the backbone network to increase the network depth at a lower cost and extract richer feature information.

[0090] Among them, dysample is a dynamic upsampling mechanism that can dynamically adjust the sampling strategy according to the size and distribution of objects in the input image, so as to adapt to more complex and diverse application environments. Dysample has two working branches. In the first branch, the input feature map passes through a point sampling generator to obtain a sampling set of sH×sW×2g, and then is input into the bilinear interpolation grid sampling together with the identity mapping of the second branch.

[0091] ​​The Coordinate Attention mechanism CA requires fewer computational resources while retaining position information. Apply the Coordinate Attention mechanism CA to YOLOv8. By embedding position information into channel attention, the mobile network can focus on larger regions while avoiding significant computational overhead. Using two one-dimensional global pooling operations, the input features in the vertical and horizontal directions are respectively aggregated into two independent direction-aware feature maps. These two feature maps embedded with specific direction information are respectively encoded into two attention maps, and each attention map captures the long-range dependencies of the input feature map along one spatial direction. Then, the two attention maps are applied to the input feature map through multiplication.

[0092] The FasterNetBlock based on Partial Convolution Pconv has the characteristics of simple structure, small number of parameters, and fast operation speed. PConv performs convolution operations on some continuous features in the input channels, and the remaining features are processed through identity mapping to keep the channels unchanged, increasing the efficiency of floating-point operation by reducing the memory access frequency of floating-point operations.

[0093] Step S4: Based on the pose information of the UAV's pan-tilt and the target detection result, perform visual target localization at the fire scene to complete the assisted fire extinguishing based on UAV information and image processing technology. When recording information, considering data processing delay, it is necessary to record when the UAV hovers stably, and the recorded information needs to be different from the previously recorded ones. Use a time window queue method to discriminate and record information, as follows:

[0094] (1) Create a queue data structure Q that can accommodate 5 groups of information. When the UAV takes off and starts to execute tasks, every 0.3 seconds, a group of pose information of the UAV's pan-tilt is enqueued. When the queue is full, dequeue first and then enqueue.

[0095] (2) Create a time series T. Every 1.5 seconds, calculate the difference degree of the 5 groups of information in Q. If the difference degree is greater than a certain threshold, it is determined that the aircraft is flying and T is assigned 1; otherwise, it is determined that the aircraft is hovering and T is assigned 0;

[0096] (3) For the time series T, when its value changes from 1 to 0, record the pose information of the UAV's pan-tilt at the moment when the value is 0.

[0097] The visual target localization method is as follows:

[0098] (1) According to the camera calibration method, obtain the camera internal parameter matrix Perform de-distortion processing on the camera.

[0099] (2) Set the origin of the camera coordinate system at the camera optical center, then the points on the imaging plane have a unified z c , and its value is the camera focal length f. According to the formula Map the output results u and v of object detection to the camera coordinate system to obtain X c =[x c y c z c Τ .

[0100] (3) Obtain the rotation matrix based on the attitude information of the UAV camera gimbal. The attitude of the UAV camera gimbal can be described by three variables: yaw angle Yaw (positive to the right and negative to the left relative to the due north direction), roll angle Roll, and pitch angle Pitch. According to these three variables, the rotation matrix R = R x R y R z , where In the formula, α is the value of Yaw, β is the value of Roll, and θ is the value of Pitch.

[0101] (4) Obtain the translation vector based on the UAV position information. Let the origin position of the world coordinate system be longitude lon0, latitude lat0, with the z-axis vertically downward as positive. The position returned by the UAV sensor is longitude lon1, latitude lat1, and height H. Calculate the difference in longitude and latitude between the UAV and the origin and convert it to radians Δlat and Δlon. Then, within a small range (within 10 - 50 kilometers), the coordinates of the UAV in the world coordinate system can be obtained. The value in the x direction (east-west direction, east is positive) is x u = r × Δlon × cos(lat0), and the value in the y direction (north-south direction, north is positive) is y u = r × Δlat. In the formula, r is the radius of the earth. Further, the translation vector L u =[x u y u -H] Τ .

[0102] (5) Based on the rotation matrix and translation vector, and according to the camera imaging principle, for a set of UAV sensor information, establish the equation R(X w -L u ) = X c to obtain X w , and realize the conversion of the pixel coordinate system - camera coordinate system - world coordinate system to obtain the coordinates of the image of the target to be located in the imaging plane in the world coordinate system.

[0103] (6) In the world coordinate system, assume that in the i-th group of UAV sensor information, the coordinates of the optical center point of the UAV camera in the world coordinate system are L o =[x 0i y 0i z 0i Τ , and pass through the optical center point of the camera and X​​w Construct a straight line For each set of UAV sensor information, a straight line can be obtained. According to the imaging model, the visual three-dimensional positioning of the target can be modeled as the following problem: In three-dimensional space, given n straight lines of the above type, find a point that minimizes the sum of the distances from this point to each straight line. Through analytic geometry and the least squares method, an approximate point can be found that minimizes the distance difference to each straight line. Let the position of the target point P in the world coordinate system be X = [x y z] Τ , the following least squares optimization problem can be established That is, minimize the sum of the squares of the distances from point P to each straight line. According to the distance formula from a point to a straight line in three-dimensional space, where:

[0104]

[0105] After arranging the above formula, we get:

[0106]

[0107] In the formula, Α1, represents the matrix composed of a set of UAV sensor information, and its calculation formula is Using the least squares method to solve, the initial solution X0 of the coordinates of the target point P to be located in the world coordinate system can be obtained.

[0108] (7) Perform non-linear optimization through bundle adjustment. For each pose observation, according to x(X) = K[R - RL u X, the pixel coordinate point x is obtained from the world coordinate point X. For the i-th image, the reprojection error is defined as e i = p i - x(X, R, L u ), where p i = [u, v] T , which is the pixel position of the target in the i-th image. To improve the three-dimensional positioning accuracy, it is necessary to optimize the reprojection error function to make it as small as possible. That is Using the pose information of the pan-tilt camera obtained by the UAV, joint constraints can also be added to the optimization target. Then the total optimization function becomes where represents the inverse of the covariance matrix. For UAVs with high positioning accuracy (such as those using RTK technology), λ can be 0. Then use non-linear least squares algorithms such as the Levenberg-Marquardt method for optimization to solve the position X of the target to be located in the world coordinate system.

[0109] As can be seen from the above mathematical derivation, multiple sets of UAV information with different poses are required for visual target positioning. To reduce the Geometric Dilution of Precision (GDOP), the UAV should make full use of its own mobility. During observation, long baselines and large intersection angles should be formed in the vertical and horizontal directions from multiple altitude layers and multiple perspectives to increase the parallax and improve the positioning accuracy.

[0110] Considering the data processing delay, it is necessary to record when the UAV hovers stably. The ground terminal program uses the method of a time window queue to automatically capture the moment when the UAV hovers stably and record the information, as follows:

[0111] (1) Create a queue data structure Q that can accommodate 5 sets of information. When the UAV takes off and starts to execute the task, every 0.3 seconds, a set of pose information of the UAV gimbal is enqueued. When the queue is full, dequeue first and then enqueue.

[0112] (2) Create a time series T. Every 1.5 seconds, calculate the difference degree of the 5 sets of information in Q. If the difference degree is greater than a certain threshold, it is determined that the aircraft is flying, and assign 1 to T; otherwise, it is determined that the aircraft is hovering, and assign 0 to T;

[0113] (3) For the time series T, when its value changes from 1 to 0, record the pose information of the UAV gimbal at the moment when the value is 0.

[0114] Step S5: Based on the target detection result and the visual target positioning, assist in dispatching fire-fighting resources to complete the fire extinguishing assistance based on UAV information and image processing technology.

[0115] This embodiment is applied to assist in dispatching fire-fighting resources, and the specific steps are as follows:

[0116] (1) The UAV acquires an image, and the image should contain a fire cannon and a fire source.

[0117] (2) Use the target detection algorithm described in step S3 to analyze the image to obtain the pose of the ground fire cannon (the orientation vector parallel to the ground), and use the visual positioning algorithm described in step S4 to obtain the position of the fire source point.

[0118] (3) In the world coordinate system, make a vector starting from the position of the fire cannon and ending at the position of the fire source point Decompose the vector into two vectors. One vector is perpendicular to the ground (i.e., parallel to the z-axis), and one vector is parallel to the ground.

[0119] (4) Calculate the vectors and the vector The included angle is the horizontal angle adjusted by the dispatching fire cannon. The vector The modulus of is the vertical distance from the fire cannon to the fire source point. According to this distance and combined with the jet mathematical model, the pitch angle can be adjusted.

[0120] Embodiment 2

[0121] An auxiliary fire extinguishing system based on unmanned aerial vehicle (UAV) information and image processing technology, the system includes:

[0122] An image acquisition module, configured to acquire fire scene images using a UAV equipped with preset equipment;

[0123] An image preprocessing module, configured to perform de-smoke preprocessing on the fire scene images to obtain preprocessed images;

[0124] A target detection module, configured to perform target detection on the preprocessed images to obtain target detection results;

[0125] A target positioning module, configured to perform visual target positioning at the fire site based on the pan-tilt pose information of the UAV and the target detection results;

[0126] A fire extinguishing assistance module, configured to assist in dispatching fire resources based on the target detection results and the visual target positioning to complete the auxiliary fire extinguishing based on UAV information and image processing technology.

[0127] In the image preprocessing module, the working process of the image preprocessing module specifically includes:

[0128] Generating a smoke image based on Perlin noise and an improved atmospheric scattering model;

[0129] Constructing a neural network model based on the UNet network and the ResNet network;

[0130] Inputting the smoke image and existing smoke image data into the neural network model to obtain model parameters;

[0131] Performing image dehazing based on the model parameters and the classical atmospheric scattering model to obtain preprocessed images.

[0132] The process of constructing a neural network model based on the UNet network and the ResNet network specifically includes:

[0133] The neural network model includes an encoder network and a decoder network;

[0134] The encoder network uses a ResNet network to extract features. The ResNet network contains 5 neural network layers. The first layer consists of a convolutional layer, a batch normalization layer, a ReLU activation function layer, and a max pooling layer in sequence. The input image is resized to a dimension of 640×640×3, and a feature map of 320×320×3 is output. The second, third, and fourth layers are residual network layers, and the residual network layer consists of a convolutional layer and a residual connection;

[0135] The decoder network consists of 5 neural network layers. The first layer receives the output of the encoder. The first, second, third, and fourth layers consist of a transposed convolutional layer, a channel concatenation layer, and a residual network layer. The input of each layer first passes through a transposed convolution, and then passes through the channel concatenation layer to be concatenated with the feature map of the same dimension size in the encoder network. Finally, it passes through the residual network layer to extract features and output. The fifth layer consists of two convolutional layers, receives the feature map of 640×640×32 output by the fourth layer of the decoder network, and the output dimension size is 640×640×3 to obtain the decoder output result.

[0136] The working process of the target positioning module specifically includes:

[0137] According to the camera calibration method, the camera internal parameter matrix is obtained, and the camera is distorted based on the camera internal parameter matrix;

[0138] The origin of the orthodontic camera coordinate system is set at the camera optical center to obtain the camera focal length. Based on the focal length and the output result of the target detection, the attitude information of the UAV camera gimbal is obtained;

[0139] Calculate the rotation matrix and translation vector based on the UAV camera gimbal attitude information;

[0140] Based on the rotation matrix, the translation vector, and the camera imaging principle, the conversion of the pixel coordinate system - camera coordinate system - world coordinate system is realized, and the coordinates of the image of the target to be located in the imaging plane under the world coordinate are obtained;

[0141] Use bundle adjustment to calculate the observations of each pose for the coordinates of the image of the target to be located in the imaging plane under the world coordinate, and obtain the visual target positioning of the fire scene based on the observation results.

[0142] The above-described embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. An auxiliary fire extinguishing method based on drone information and image processing technology, characterized in that, The method includes: Step S1: Use a drone equipped with preset equipment to collect fire scene images; Step S2: Perform de-smoking preprocessing on the fire scene images to obtain preprocessed images; Step S3: Perform target detection on the preprocessed images to obtain target detection results; Step S4: Based on the pan-tilt pose information of the drone and the target detection results, perform visual target positioning at the fire scene; Step S5: Based on the target detection results and the visual target positioning, assist in dispatching fire resources to complete fire extinguishing assistance based on drone information and image processing technology.

2. The auxiliary fire extinguishing method based on drone information and image processing technology according to claim 1, wherein In step S2, the process of obtaining the preprocessed images specifically includes: Construct a neural network model based on the UNet network and the ResNet network. During neural network training, use the Berlin noise to generate model parameters t(x) and c(x), input the model parameters t(x) and c(x) into the improved atmospheric scattering model to generate a smoke image, and use the smoke image and the existing fire scene images to train the neural network model to obtain a trained neural network model; Input the fire scene images into the trained neural network model to obtain preprocessed images.

3. The auxiliary fire extinguishing method based on drone information and image processing technology according to claim 2, wherein, The process of constructing the neural network model based on the UNet network and the ResNet network specifically includes: The neural network model includes an encoder network and a decoder network; The encoder network uses the ResNet network to extract features. The ResNet network contains 5 neural network layers. The first layer consists of a convolutional layer, a batch normalization layer, a ReLU activation function layer, and a max pooling layer in sequence. The input image is resized to a dimension size of 640×640×3, and a feature map of 320×320×3 is output. The second, third, and fourth layers are residual network layers, and the residual network layer consists of a convolutional layer and a residual connection; The decoder network consists of 5 neural network layers. The first layer receives the output of the encoder. The first, second, third, and fourth layers consist of a transposed convolutional layer, a channel concatenation layer, and a residual network layer. The input of each layer first passes through a transposed convolution, and then passes through the channel concatenation layer to be concatenated with the feature map of the same dimension size in the encoder network, and finally passes through the residual network layer to extract features and output. The fifth layer consists of two convolutional layers, receives the feature map of 640×640×32 output by the fourth layer of the decoder network, and outputs a dimension size of 640×640×3 to obtain the decoder output result.

4. The auxiliary fire extinguishing method based on drone information and image processing technology according to claim 3, wherein, The process of using the Berlin noise to generate model parameters t(x) and c(x) and inputting the model parameters t(x) and c(x) into the improved atmospheric scattering model to generate a smoke image specifically includes: The generation process of t(x) specifically includes: Set the number of noise layers to 6, persistence to 0.5, porosity to 2.0, scale to 250, and generate the Perlin noise map P after normalization noise ; Based on the Berlin noise map, generate a smoke transmittance map t(x); t(x) = P noise ·(t max - t min ) + t min where t max is the maximum value of the transmittance, and t min is the minimum value of the transmittance; The generation process of c(x) specifically includes: Based on the Berlin noise map, generate the RGB value c(x) of the smoke haze; c(x) = P noise ·(c max -c min ) + c min ; Among them, c max is the maximum value of the RGB of the smoke and haze, and c max is the minimum value of the RGB of the smoke and haze; Substitute the obtained t(x) and c(x) into the improved atmospheric scattering model to add smoke to the clear image to obtain a smoke image.

5. The auxiliary fire extinguishing method based on drone information and image processing technology according to claim 4, wherein The improved atmospheric scattering model is specifically: I(x) = J(x)t(x) + c(x)(1 - t(x)) Where, I(x) is the fire scene image; J(x) is the image after ideal haze removal processing; t(x) is the smoke transmittance map; c(x) is the RGB value of the smoke and haze.

6. The auxiliary fire extinguishing method based on drone information and image processing technology according to claim 1, characterized in that, In the step S4, the process of visually locating the target at the fire scene specifically includes: According to the camera calibration method, obtain the camera internal parameter matrix, and perform distortion processing on the camera based on the camera internal parameter matrix; Use the orthodontic UAV positioning module to obtain the position information of the UAV; Obtain the attitude information of the UAV gimbal camera based on the orthodontic UAV sensor; Obtain the translation vector based on the position information of the UAV, calculate the rotation matrix based on the attitude information of the UAV camera gimbal, and obtain the imaging external parameter matrix; Based on the target detection result, the imaging external parameter matrix, and the camera imaging principle, realize the conversion of the pixel coordinate system - camera coordinate system - world coordinate system, and obtain the initial coordinate solution of the target to be located in the imaging plane under the world coordinate; Use bundle adjustment to optimize the three-dimensional coordinates of the target to be located according to the observation results of each pose of the UAV, and obtain the visual target positioning at the fire scene.

7. An auxiliary fire extinguishing system based on drone information and image processing technology, the system is used to implement the positioning method described in any one of claims 1-6, characterized in that, The system includes: An image acquisition module for using a UAV equipped with preset equipment to acquire fire scene images; An image preprocessing module for performing de-smoke preprocessing on the fire scene image to obtain a preprocessed image; A target detection module for performing target detection on the preprocessed image to obtain a target detection result; A target positioning module for performing visual target positioning at the fire scene based on the gimbal pose information of the UAV and the target detection result; A fire extinguishing assistance module for assisting in dispatching fire resources based on the target detection result and the visual target positioning, and completing the assisted fire extinguishing based on UAV information and image processing technology.

8. The auxiliary fire extinguishing system based on drone information and image processing technology according to claim 7, characterized in that, The process of obtaining the preprocessed image specifically includes: Construct a neural network model based on the UNet network and the ResNet network. During the neural network training, use the Perlin noise generation model to generate the model parameters t(x) and c(x), input the model parameters t(x) and c(x) into the improved atmospheric scattering model to generate a smoke image, and use the smoke image and the existing fire scene image to train the neural network model to obtain a trained neural network model; Input the fire scene image into the trained neural network model to obtain a preprocessed image.

9. The auxiliary fire extinguishing system based on drone information and image processing technology according to claim 8, wherein The process of constructing the neural network model based on the UNet network and the ResNet network specifically includes: The neural network model includes an encoder network and a decoder network; The encoder network uses the ResNet network to extract features. The ResNet network contains 5 neural network layers. The first layer is composed of a convolutional layer, a batch normalization layer, a ReLU activation function layer, and a max pooling layer in sequence. The input image is Resized to a dimension size of 640×640×3, and a feature map of 320×320×3 is output. The second layer, the third layer, and the fourth layer are residual network layers, and the residual network layer is composed of a convolutional layer and a residual connection; The decoder network consists of 5 neural network layers. The first layer receives the output of the encoder. The first, second, third, and fourth layers are composed of transposed convolutional layers, channel concatenation layers, and residual network layers. The input of each layer first passes through a transposed convolution, then through a channel concatenation layer to concatenate channels with the feature map of the same dimension size in the encoder network, and finally through a residual network layer to extract features and output. The fifth layer is composed of two convolutional layers, receives the 640×640×32 feature map output by the fourth layer of the decoder network, and outputs a dimension size of 640×640×3 to obtain the decoder output result.

10. The auxiliary fire extinguishing system based on drone information and image processing technology according to claim 8, characterized in that, The process of using the Berlin noise generation model parameters t(x) and c(x) and inputting the model parameters t(x) and c(x) into the improved atmospheric scattering model to generate a smoke image specifically includes: The generation process of the t(x) specifically includes: Set the number of noise layers to 6, the persistence to 0.5, the porosity to 2.0, the scale to 250, and generate the Perlin noise map P after normalization noise ; Based on the Berlin noise map, generate a smoke transmittance map t(x): t(x) = P noise ·(t max -t min ) + t min where t max is the maximum value of the transmittance, and t min is the minimum value of the transmittance; The generation process of the c(x) specifically includes: Based on the Berlin noise map, generate the RGB value c(x) of the smoke haze: c(x) = P noise ·(c max -c min ) + c min ; Among them, c max is the maximum value of the RGB of the smoke and haze, and c max is the minimum value of the RGB of the smoke and haze; Substitute the obtained t(x) and c(x) into the improved atmospheric scattering model to add smoke to the clear image to obtain a smoke image.

Citation Information

Patent Citations

  • Simultaneous positioning and mapping method for autonomous mobile platform in rescue scene

    CN111583136A

  • Image defogging method and system and embedded device

    CN116228580A

  • Unmanned aerial vehicle visual fire extinguishing decision-making method and system

    CN117218514A

  • Video defogging method and system based on double constraints of color domain and frequency domain

    CN117291828A

  • Laparoscopic image smoke removal method based on generative adversarial network

    US20230377097A1