Multi-target fire positioning method and device based on multi-spectral dynamic fusion
By combining visible light and infrared images with a multispectral dynamic fusion method to identify smoke and fire sources, and using drone posture data for ray inversion, the problems of high false detection rate and low efficiency of multi-target positioning in drone fire detection are solved, and efficient and accurate multi-target fire positioning is achieved.
Patent Information
- Application Number
- CN202511022685.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-07-24
AI Technical Summary
Existing drone fire detection technology relies on a single sensor, resulting in a high false detection rate and poor adaptability. In addition, the image is blurred and the target is offset when flying on the inspection route, making it difficult to simultaneously process multiple fire targets and resulting in low positioning efficiency.
A multispectral dynamic fusion method is used to combine visible light and infrared images to identify smoke and fire sources. The target area is determined through confidence fusion processing, and ray inversion is performed using drone posture data to achieve precise positioning of multi-target fires.
The efficiency and accuracy of UAV identification and positioning of multi-target fire areas during patrol missions have been improved, and the accuracy and robustness of positioning have been enhanced.
Smart Images

Figure CN120526237B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of unmanned aerial vehicle inspection technology, and in particular to a multi-target fire positioning method and device based on multi-spectral dynamic fusion. Background Art
[0002] With the rapid development of drone technology, its application in fields such as forest fire prevention and urban firefighting is becoming increasingly widespread. Currently, existing technologies typically use only a single sensor type (such as infrared or visible light sensors) for detection. However, in real-world scenarios, interference sources such as high-temperature objects and sunlight reflections have similar characteristics to actual fire sources. Smoky environments can also affect the accuracy of visible light sensors, resulting in a high false positive rate. Furthermore, when drones are in patrol mode, factors such as body vibration and heading changes can blur captured images and cause target offsets, leading to significant deviations between positioning results and the actual fire source location. Furthermore, in complex fire environments, existing technologies struggle to simultaneously process multiple targets while the drone is in patrol mode, resulting in low positioning efficiency. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide a multi-target fire positioning method and device based on multi-spectral dynamic fusion, which can not only realize the identification and positioning of multiple target areas during the UAV's cruise mission, but also help improve the positioning efficiency of the target area, and can also significantly improve the positioning accuracy of the target area.
[0004] In a first aspect, the present invention provides a multi-target fire location method based on multi-spectral dynamic fusion, comprising:
[0005] Acquire a multispectral image sequence collected by the UAV during an inspection mission. The multispectral image sequence includes multiple image frame groups, each of which includes a visible light image and an infrared image at the same timestamp.
[0006] Performing smoke recognition processing, fire source recognition processing, and confidence-based dynamic fusion processing on the visible light image and the infrared image included in the image frame group to obtain target pixel positions identified as target areas in the image frame group. The target area is an area where the results of the smoke recognition processing and the results of the fire source recognition processing are both located. There is at least one target area.
[0007] Based on the target pixel position identified as the target area in the image frame group, ray inversion is performed using the drone's posture data as positioning compensation to determine the target latitude and longitude positions of the target area identified as the target area in the image frame group.
[0008] In one embodiment, smoke recognition processing, fire source recognition processing, and confidence-based dynamic fusion processing are performed based on the visible light image and infrared image included in the image frame group to obtain the target pixel position identified as the target area in the image frame group, including:
[0009] Performing smoke recognition processing on the visible light image at the current timestamp to obtain a smoke mask region, determining the smoke area and mask edge sharpness corresponding to the smoke mask region, and using this to obtain a smoke confidence level corresponding to the smoke mask region;
[0010] Performing fire source identification processing on the infrared image at the current timestamp to obtain a fire source mask area, determining the thermal feature change amplitude and mask stability corresponding to the fire source mask area, and obtaining the fire source confidence corresponding to the fire source mask area;
[0011] Based on the smoke confidence and the fire source confidence, the pixel position of the specified point in the smoke mask area and the pixel position of the specified point in the fire source mask area are fused to obtain the target pixel position identified as the target area at the current timestamp.
[0012] In one embodiment, determining the thermal signature variation amplitude and mask stability corresponding to the fire source mask area to obtain the fire source confidence level corresponding to the fire source mask area includes:
[0013] Extract the highest and lowest temperature values within the fire source mask area at the current timestamp, and determine the thermal feature change amplitude corresponding to the fire source mask area of the current frame in combination with the standard temperature difference reference value;
[0014] and, determining the mask stability corresponding to the fire source mask area at the current timestamp based on the center of gravity change distance and mask area change ratio between the fire source mask area at the current timestamp and the fire source mask area at the previous timestamp;
[0015] The thermal feature variation amplitude and mask stability are weightedly fused to obtain the fire source confidence corresponding to the fire source mask area at the current timestamp.
[0016] In one embodiment, based on the target pixel position identified as the target area in the image frame group, ray inversion is performed using the drone's posture data as positioning compensation to determine the target latitude and longitude position of the target area identified as the target area in the image frame group, including:
[0017] Based on the target pixel position identified as the target area at the current timestamp, and the motion parameters and gimbal angle contained in the drone's posture data, an observation ray vector pointing from the drone to the target area at the current timestamp is constructed;
[0018] Determine the observation weight values corresponding to multiple target timestamps, perform ray inversion on the current timestamp based on the observation ray vector and observation weight value at the target timestamp, and obtain the target latitude and longitude positions identified as the target area at the current timestamp.
[0019] In one embodiment, based on the target pixel position identified as the target area at the current timestamp and the motion parameters and gimbal angle included in the drone's posture data, constructing an observation ray vector pointing from the drone to the target area at the current timestamp includes:
[0020] Based on the UAV's motion parameters, the UAV's position at the current timestamp is calculated;
[0021] According to the gimbal angle of the drone and the target pixel position identified as the target area at the current timestamp, the ray direction vector corresponding to the target pixel position at the current timestamp is determined;
[0022] Based on the ray direction vectors corresponding to the drone position and the target pixel position, the observation ray vector pointing from the drone to the target area at the current timestamp is constructed.
[0023] In one embodiment, determining observation weight values corresponding to a plurality of target timestamps includes:
[0024] For any target timestamp, perform the following operations:
[0025] Based on the corresponding values of the image clarity item, segmentation confidence item, attitude stability item and triangulation baseline perspective rationality item at the target timestamp, the observation weight value corresponding to the target timestamp is determined;
[0026] Among them, the image clarity item is used to describe the influence of the variance of the image recognition results of visible light images and infrared images on the observation weight value; the segmentation confidence item is used to describe the influence of the smoke confidence and fire source confidence on the observation weight value; the attitude stability item is used to describe the influence of the degree of change of the attitude angular velocity of the UAV on the observation weight value; the reasonable triangulation baseline perspective item is used to describe the influence of the angle between the UAV's perspective and the horizontal line and the length of the UAV's triangulation baseline on the observation weight value.
[0027] In one embodiment, ray inversion is performed based on the observation ray vector and the observation weight value at the target timestamp to obtain the target latitude and longitude position identified as the target area at the current timestamp, including:
[0028] By using a pre-built objective function, based on the observation ray vector and the observation weight value at the target timestamp, the target ground position identified as the target area at the current timestamp is determined so that the distance between the target ground position and the vertical position of the observation ray vector at the target timestamp is the shortest;
[0029] The target ground position is converted into coordinates to obtain the target latitude and longitude position identified as the target area at the current timestamp.
[0030] In one embodiment, the method further comprises:
[0031] The target longitude and latitude positions of the same target area at different timestamps are spliced together to obtain the target trajectory sequence corresponding to each target area. For any target trajectory sequence, perform at least one of the following optimization processes:
[0032] An observation model is constructed based on the target latitude and longitude positions contained in the target trajectory sequence, and the target trajectory sequence is optimized by Kalman filtering using the observation model to smooth the target trajectory sequence;
[0033] For any timestamp in the target trajectory sequence, if the observation weight value corresponding to the timestamp is lower than the preset weight threshold, the target latitude and longitude position at the timestamp is removed from the target trajectory sequence;
[0034] For any timestamp in the target trajectory sequence, determine the spatial variation distance between the target latitude and longitude position at that timestamp and the target latitude and longitude position at the previous timestamp. If the spatial variation distance is greater than a preset distance threshold, remove the target latitude and longitude position at that timestamp from the target trajectory sequence.
[0035] For any timestamp in the target trajectory sequence, a regional consistency check is performed on the target pixel position and the target latitude and longitude position at that timestamp. If the regional consistency check fails, the target latitude and longitude position at that timestamp is removed from the target trajectory sequence.
[0036] In a second aspect, the present invention further provides a multi-target fire location device based on multi-spectral dynamic fusion, comprising:
[0037] An image acquisition module is used to acquire a multispectral image sequence collected by the UAV during the inspection mission. The multispectral image sequence includes multiple image frame groups, and the image frame group includes visible light images and infrared images at the same time stamp;
[0038] an identification and fusion module for performing smoke identification processing, fire source identification processing, and dynamic fusion processing of the results based on confidence levels based on the visible light image and infrared image contained in the image frame group, to obtain the target pixel position identified as the target area in the image frame group. The target area is the area where the results of the smoke identification processing and the results of the fire source identification processing are located together, and there is at least one target area;
[0039] The position inversion module is used to perform ray inversion based on the target pixel position identified as the target area in the image frame group, using the posture data of the drone as positioning compensation, to determine the target latitude and longitude position of the target area identified as the target area in the image frame group.
[0040] In a third aspect, the present invention further provides an electronic device comprising a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement any one of the methods provided in the first aspect.
[0041] The present invention provides a multi-target fire positioning method and device based on multispectral dynamic fusion. First, a multispectral image sequence collected by a drone during an inspection mission is obtained, where the multispectral image sequence includes multiple image frame groups, and the image frame group includes visible light images and infrared images at the same timestamp; then, smoke recognition processing, fire source recognition processing, and dynamic fusion processing of the results based on confidence are performed based on the visible light images and infrared images contained in the image frame group to obtain the target pixel position identified as the target area in the image frame group, where the target area is the area where the results of the smoke recognition processing and the results of the fire source recognition processing are located together, and there is at least one target area; finally, based on the target pixel position identified as the target area in the image frame group, ray inversion is performed using the drone's posture data as positioning compensation to determine the target latitude and longitude position of the target area identified as the target area in the image frame group. The above method uses visible light images and infrared images to perform smoke recognition, fire source recognition and dynamic fusion of confidence-based results to obtain at least one target pixel position identified as a target area. On this basis, ray inversion is performed using the drone's posture data as positioning compensation to achieve precise positioning of multiple target areas and obtain the target longitude and latitude positions. The present invention can not only realize the identification and positioning of multiple target areas during the drone's cruise mission, which helps to improve the positioning efficiency of the target area, but also significantly improve the positioning accuracy of the target area.
[0042] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or understood by practicing the present invention. The purposes and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.
[0043] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0045] Figure 1 A schematic flow chart of a multi-target fire location method based on multi-spectral dynamic fusion provided by an embodiment of the present invention;
[0046] Figure 2 A schematic diagram of the overall process of a multi-target fire location method based on multi-spectral dynamic fusion provided by an embodiment of the present invention;
[0047] Figure 3 A schematic structural diagram of a multi-target fire location device based on multi-spectral dynamic fusion provided by an embodiment of the present invention;
[0048] Figure 4 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0050] At present, the existing technology has the following problems: traditional drone fire source detection relies on a single sensor, which makes it difficult to effectively distinguish between fire sources and smoke, and has problems such as high false detection rate and poor adaptability; when the drone is in the inspection route flight state, the drone movement causes image jitter and target offset, and the traditional ray method does not introduce motion parameter compensation, which affects positioning accuracy; traditional drone fire source detection has difficulty in synchronously processing multiple targets under dynamic conditions, resulting in a lack of efficient data fusion and target solution mechanism.
[0051] Based on this, the present invention provides a multi-target fire positioning method and device based on multi-spectral dynamic fusion, which can not only realize the identification and positioning of multiple target areas during the UAV's cruise mission, helping to improve the positioning efficiency of the target area, but also significantly improve the positioning accuracy of the target area.
[0052] To facilitate understanding of this embodiment, firstly, a multi-target fire location method based on multi-spectral dynamic fusion disclosed in an embodiment of the present invention is described in detail. Figure 1 The flowchart of a multi-target fire location method based on multi-spectral dynamic fusion is shown, and the method mainly includes the following steps S102 to S106:
[0053] Step S102: Acquire a multispectral image sequence collected by the UAV during the inspection mission.
[0054] The multispectral image sequence includes multiple image frame groups, each of which includes visible light images and infrared images at the same timestamp. In one embodiment, a drone equipped with a gimbal performs inspections along a route, capturing visible light and infrared images during the inspection process and uploading them to a server in real time. This captures visible light and infrared images at multiple timestamps, forming a multispectral image sequence.
[0055] Step S104 , performing smoke recognition processing, fire source recognition processing, and confidence-based dynamic fusion processing on the visible light image and infrared image included in the image frame group, to obtain the target pixel position identified as the target area in the image frame group.
[0056] The target area is the area where the results of the smoke identification process and the fire source identification process coexist, that is, the area where the smoke identified based on the visible light image overlaps with the fire source identified based on the infrared image. There is at least one target area, and the target pixel position is the pixel position corresponding to the fire source point (smoke point) located in the target area. In one embodiment, smoke identification is performed using the visible light image to obtain a corresponding smoke mask area and a smoke confidence level. The smoke confidence level is determined based on the smoke area and mask edge sharpness corresponding to the smoke mask area. Fire source identification is performed using the infrared image to obtain a corresponding fire source mask area and a fire source confidence level. The fire source confidence level is determined based on the thermal signature variation amplitude and mask stability corresponding to the fire source mask area. Based on the smoke confidence level and the fire source confidence level, the center pixel positions of the smoke mask area and the fire source mask area are fused to obtain the target pixel position identified as the target area.
[0057] Step S106 , based on the target pixel position identified as the target area in the image frame group, ray inversion is performed using the drone's posture data as positioning compensation to determine the target latitude and longitude position of the target area identified as the target area in the image frame group.
[0058] The pose data includes motion parameters, camera parameters, and gimbal angles. Motion parameters may include motion speed, camera parameters may include field of view angle, resolution, camera principal point coordinates, and camera focal length, and gimbal angles may include pitch angle, roll angle, and yaw angle. The target latitude and longitude positions are the coordinates of the target area in the WGS84 coordinate system. In one embodiment, the motion parameters are used to infer the drone position at the current timestamp, and the ray direction vector corresponding to the target pixel position is determined based on the target pixel position identified as the target area at the current timestamp. The drone position and ray direction vector are combined to construct an observation ray vector pointing to the target area at the current timestamp. The observation ray vectors at multiple timestamps are used to solve the target latitude and longitude positions identified as the target area at the current timestamp.
[0059] The multi-target fire positioning method based on multi-spectral dynamic fusion provided by the embodiment of the present invention performs smoke recognition, fire source recognition and dynamic fusion of confidence-based results through visible light images and infrared images to obtain the target pixel position of at least one target area identified as the target area. On this basis, ray inversion is performed using the posture data of the drone as positioning compensation to achieve precise positioning of multiple target areas and obtain the target longitude and latitude positions. The present invention can not only realize the identification and positioning of multiple target areas during the drone's cruise mission, which helps to improve the positioning efficiency of the target area, but also significantly improve the positioning accuracy of the target area.
[0060] For ease of understanding, the present invention provides a specific implementation of a multi-target fire location method based on multi-spectral dynamic fusion, see Figure 2 The overall process diagram of a multi-target fire location method based on multi-spectral dynamic fusion is shown, including the following steps S202 to S214:
[0061] Step S202: Acquire a multispectral image sequence collected by the UAV during the inspection mission.
[0062] Step S204 , performing smoke recognition processing on the visible light image at the current timestamp to obtain a smoke mask region, determining the smoke area and mask edge sharpness corresponding to the smoke mask region, and obtaining a smoke confidence corresponding to the smoke mask region.
[0063] In one example, smoke recognition can be performed using a model based on an improved U-Net structure. The input of the model is a visible light image, and the output is a smoke mask area. .
[0064] In one example, the process of determining smoke confidence is as follows:
[0065] Count the visible light image segmented into smoke mask areas Number of pixels , count the total number of pixels in the visible light image , number of pixels Total number of pixels It reflects whether the smoke identified in the visible light image at the current timestamp is macroscopically observable;
[0066] Mask area for smoke Perform edge detection. The edge detection algorithm can use Sobel or Canny, etc., and count the smoke mask area. The gradient intensity of the edge pixels in the middle is the gradient amplitude of Sobel x / y. The average gradient intensity of all edge pixels (ie, the average gradient intensity) is used as the smoke mask area. EdgeSharpness of the mask edge. EdgeSharpness is an image quality indicator that measures the clarity of the edge contour of the smoke mask area. It is usually used to reflect the stability or reliability of the identified smoke mask area.
[0067] Number of pixels Total number of pixels The smoke confidence is obtained by weighted fusion of the ratio of the mask edge sharpness EdgeSharpness .
[0068] Specifically, smoke confidence The calculation formula is as follows:
[0069] ;
[0070] in, is the smoke confidence, The visible light image is segmented into the smoke mask area The number of pixels, is the total number of pixels in the visible light image, Mask area for smoke The mask edge sharpness, 、 is a weighting coefficient, which is used to adjust the contribution of the two factors "smoke area" and "mask edge sharpness" to the confidence level. Both can be set through training set regression fitting, network search or experience parameter adjustment. Specifically, Control the influence of the proportion of smoke area in the image on the smoke confidence. Controls the effect of edge sharpness on smoke confidence.
[0071] Step S206 , performing fire source identification processing on the infrared image at the current timestamp to obtain a fire source mask area, determining the thermal feature variation amplitude and mask stability corresponding to the fire source mask area, and obtaining the fire source confidence corresponding to the fire source mask area.
[0072] In one example, fire source identification can be performed using a ResNet-based thermal imaging semantic segmentation network. The input of the thermal imaging semantic segmentation network is an infrared image, and the output is a fire source mask area. .
[0073] In one example, the process of determining the confidence level of a fire source is as follows:
[0074] Extract the fire source mask area at the current timestamp The maximum temperature value within and the minimum temperature , combined with the standard temperature difference reference value Determine the fire source mask area of the current frame The corresponding thermal feature change range. Specifically: the highest temperature value in the fire source mask area can be obtained by restoring the thermal value of the original infrared image or pseudo-color image. and minimum temperature , and set an empirical value as the standard temperature difference reference value for normalization , where the highest temperature value , minimum temperature value and standard temperature difference reference value Both are used to evaluate the magnitude of thermal signature changes in the fire source mask area in infrared images.
[0075] Fire source mask area based on the current timestamp The fire source mask area at the previous time stamp The distance between the center of gravity changes and the mask area change ratio are used to determine the fire source mask area at the current timestamp. Corresponding mask stability .Mask stability Refers to the stability of the fire source mask area under two consecutive time stamps, which is used to suppress occasional false alarms caused by interference such as thermal noise and vehicle exhaust. Specifically: Compare the fire source mask area The distance of the center of gravity change : , The fire source mask area at the current timestamp The center of gravity, The fire source mask area at the previous time stamp Center of gravity; compare the fire source mask area The mask area change ratio : , The fire source mask area at the current timestamp The mask area, The fire source mask area at the previous time stamp The mask area; the mask stability is determined by combining the center of gravity change distance and the mask area change ratio : , 、 is the weighting coefficient.
[0076] Perform weighted fusion of the thermal feature change amplitude and mask stability to obtain the fire source confidence corresponding to the fire source mask area at the current timestamp .
[0077] Specifically, the fire source confidence The calculation formula is as follows:
[0078] ;
[0079] in, is the fire source confidence, is the maximum temperature value, is the minimum temperature value, is the standard temperature difference reference value, For mask stability, 、 is a weighting coefficient, which is used to adjust the contribution of the two factors "thermal feature variation amplitude" and "mask stability" to the confidence level. Both can be set through training set regression fitting, network search or experience parameter adjustment. Specifically, Control the influence of the variation of thermal characteristics on the confidence of fire source, The effect of control mask stability on fire source confidence.
[0080] In step S208, based on the smoke confidence level and the fire source confidence level, the pixel positions of the designated points in the smoke mask area and the designated points in the fire source mask area are fused to obtain the target pixel position identified as the target area at the current timestamp. The designated points in the smoke mask area and the fire source mask area can both be midpoints. Optionally, the geometric center points of the smoke mask area and the fire source mask area can be used as midpoints. In one embodiment, if the smoke mask area overlaps with the fire source mask area, a confidence-weighted average fusion is used to obtain the target pixel position. The target pixel position is calculated as follows:
[0081] ;
[0082] ;
[0083] in, , is the target pixel position for subsequent spatial calculation after fusion. , is the center pixel position of the smoke mask area, , is the midpoint pixel position of the fire source mask area, , It is the smoke confidence and fire source confidence, and its role is to use more reliable image information to guide the final positioning and identification.
[0084] These target pixel coordinates will then be converted into line of sight angles (yaw angle, pitch angle offset), and further input into the ray model for spatial positioning. The core idea that this formula wants to express is: when the smoke mask area and the fire source mask area overlap, the confidence of the two is combined to fuse their central pixel positions to obtain a more reliable target pixel position for spatial calculation. In practical applications, infrared fire sources and visible light smoke usually appear in the same area at the same time, but the response position, range and accuracy of the two sensors to the area are inconsistent: infrared images may cause positioning offsets due to heat diffusion or false fire source interference, and the smoke mask area in the visible light image may be blurred due to illumination or occlusion. Therefore, you cannot simply choose one of them, but should fuse the two. In an embodiment of the present invention, the confidence reflects the dynamic nature of the trust source: if the smoke confidence If it is very high, it means that the smoke detection is more reliable, so the fusion center is more biased towards the smoke mask area. The center pixel position of , ); If the fire source confidence Very high, indicating that the fire source detection is more reliable, so the fusion center is more inclined to the fire source mask area The center pixel position of , ).
[0085] Step S210: Based on the target pixel position identified as the target area at the current timestamp, and the motion parameters and gimbal angle contained in the drone's posture data, construct an observation ray vector pointing from the drone to the target area at the current timestamp. Specifically, it includes:
[0086] (1) Synchronize UAV motion parameters, including real-time position information (WGS84) and velocity vector , Represents the components of velocity in the x, y, and z directions; gimbal angle: pitch angle , yaw angle , roll angle .
[0087] (2) Based on the UAV's motion parameters, the UAV's position at the current timestamp is calculated. Specifically, the UAV's position at the current timestamp is calculated according to the following formula: The drone's location :
[0088] ;
[0089] ;
[0090] in, The most recently available flight control position data timestamp, represents the position change, 、 Both indicate The current timestamp corresponding to the image frame In practical applications, the current timestamp of the image frame is taken according to the image module acquisition frequency, such as 10 frames per second, but the timestamp of the flight control position data (GPS / INS) is usually updated at a lower frequency, such as 5 times per second, so the current timestamp needs to be determined by interpolation. The drone's location , drone location Used for subsequent construction of observation ray vectors = .
[0091] In this embodiment of the present invention, the position of the drone is calculated The purpose is to improve the spatial longitude of the three-dimensional ray of the image frame and ensure the spatiotemporal consistency of the solution of "image → ray → ground point".
[0092] (3) According to the gimbal angle of the UAV and the target pixel position identified as the target area at the current timestamp, determine the ray direction vector corresponding to the target pixel position at the current timestamp.
[0093] In one embodiment, the target pixel position ( , ) and camera parameters (field of view , resolution ) to obtain the offset angle, is the horizontal field of view, is the vertical field of view, is the horizontal resolution, is the vertical resolution:
[0094] ;
[0095] ;
[0096] in, is the angular offset of the target pixel position relative to the image center in the horizontal direction (yaw axis), is the angular offset of the target pixel position in the vertical direction (pitch axis), and the target pixel position ( , ) is converted into an offset angle in the direction of the optical axis 、 Used for subsequent construction of ray direction vector , ray direction vector The calculation formula is as follows:
[0097] .
[0098] The essence of the above calculation method is to directly merge the camera angle of view offset into the yaw angle and pitch angle However, this calculation method has poor accuracy. Therefore, the embodiment of the present invention provides another method for determining the ray direction vector corresponding to the target pixel position at the current timestamp. Implementation method to obtain more accurate ray direction vector :
[0099] For the current timestamp An identified target region in the image is identified based on its target pixel position ( , ), and get the corresponding ray direction vector in the camera coordinate system:
[0100] ;
[0101] in, It is The current timestamp corresponding to the image frame The target pixel position under is the ray direction vector corresponding to the camera coordinate system. , For the The current timestamp corresponding to the image frame The target pixel position under , is the coordinate of the principal point of the camera (that is, the intersection of the optical axis and the image plane), is the camera focal length in pixels.
[0102] The ray direction vector Convert from the camera coordinate system to the geographic coordinate system (ENU coordinate system) and the conversion relationship is:
[0103] ;
[0104] in, For the The current timestamp corresponding to the image frame The target pixel position under the ray direction vector corresponding to the geographic coordinate system, For the The gimbal attitude rotation matrix at the frame time (from camera system to aircraft system), For the The body posture rotation matrix at the frame time (from the body system to the ENU system), the ray direction vector Used for subsequent construction of observation ray vectors = .
[0105] Gimbal attitude rotation matrix It is a rotation matrix used to transform the target direction from the camera coordinate system to the drone body coordinate system. It can be obtained by gimbal angle information (pitch angle , yaw angle , roll angle ) is calculated. Generally speaking, the gimbal on the drone can report angle information by itself for calculation of the rotation matrix. The embodiment of the present invention uses the following formula to determine the gimbal attitude rotation matrix : ;in, 、 Represents the yaw angle , pitch angle , roll angle The matrix of .
[0106] (4) Based on the ray direction vectors corresponding to the drone position and the target pixel position, construct the observation ray vector pointing from the drone to the target area at the current timestamp. Specifically, the expression of the observation ray vector is as follows: = .
[0107] Step S212, determine the observation weight values corresponding to multiple target timestamps, perform ray inversion on the current timestamp based on the observation ray vector and observation weight value at the target timestamp, and obtain the target latitude and longitude position identified as the target area at the current timestamp.
[0108] To improve the accuracy and robustness of ground target positioning, this embodiment of the present invention uses a multi-view fusion ray least squares optimization method to estimate ground target coordinates. This embodiment utilizes the same target region in multiple image frames from different times or locations, and jointly solves for the intersection point closest to all observed ray vectors through least squares, thereby avoiding single-frame intersection errors.
[0109] Ideally, multiple observation ray vectors should intersect at the true position of the ground target, but due to the influence of noise and attitude error, they do not actually intersect at the same point. Therefore, the embodiment of the present invention searches for a point , so that it is as close as possible to all observation ray vectors:
[0110] ;
[0111] in, is the number of observation ray vectors involved in positioning in multi-view observations, is the identity matrix, that is , used to construct the projection matrix The optimization goal is to solve a point , is the minimum vertical distance to each observation ray vector, the projection matrix in the formula Project the vector in the direction perpendicular to the observation ray vector.
[0112] Another error term Defined as: ;
[0113] In traditional least squares, all observation ray vectors are assumed to be equally reliable. Has a non-negative observation weight value The larger the observation weight value, the greater the observation ray vector The more reliable it is, the greater its proportion in the solution process. In one optional implementation, a sliding window can be constructed for the current timestamp, and the timestamp contained in the sliding window can be used as the target timestamp. The observation weight value corresponding to each target timestamp is determined. The process is as follows: Based on the corresponding values of the image clarity item, segmentation confidence item, posture stability item, and triangulation baseline view reasonableness item at the target timestamp, the observation weight value corresponding to the target timestamp is determined.
[0114] In a specific implementation, the product combination function of the above indicators can be: Image clarity item It is used to describe the influence of the variance of image recognition results of visible light images and infrared images on the observation weight value. Specifically: , is the variance of the image recognition result. Segmentation confidence item Used to describe the impact of smoke confidence and fire source confidence on the observation weight value, specifically: split confidence item The target mask confidence score directly output by the neural network. It is used to describe the influence of the change of the UAV's attitude angular velocity on the observation weight value. Specifically: , is the attitude angular velocity. The more drastic the change in attitude angular velocity, the lower the weight. It is used to describe the influence of the angle between the drone's viewing angle and the horizontal line and the length of the drone's triangulation baseline on the observation weight value. Specifically: , That is, The pitch angle corresponding to the frame. When the viewing angle is too parallel and the triangulation baseline is too short, its weight can be reduced. Furthermore, all observation ray vectors can be The observation weight value Perform normalization: , That is, The observation weight value of the observation ray vector.
[0115] After determining the observation weight value, the pre-built objective function can be used to determine the target ground position identified as the target area at the current timestamp based on the observation ray vector and the observation weight value at the target timestamp, so that the vertical distance between the target ground position and the observation ray vector at all target timestamps is The shortest. The objective function is expressed as follows:
[0116] ;
[0117] in, It is the number of observation ray vectors involved in positioning in multi-view observation.
[0118] right Taking the derivative and setting it to 0, we get:
[0119] ;
[0120] because is a symmetric projection matrix, which simplifies to:
[0121] ;
[0122] in, , that is is a symmetric projection matrix;
[0123] The final target ground position is: .
[0124] Perform coordinate conversion on the target ground position to obtain the current timestamp The target latitude and longitude locations identified as the target area In the specific implementation, the target ground position The target area is located in the ENU coordinate system. The target ground position can be calculated based on the original starting point geographic information of the UAV. Convert the ENU coordinate system to the WGS84 coordinate system to obtain the target longitude and latitude position. : ,in( ) is the WGS84 coordinate of the reference origin in the ENU coordinate system, representing the three-dimensional geographic location of the detection target such as smoke or fire source. Represents the coordinate conversion matrix from the ENU coordinate system to the longitude, latitude, and altitude (WGS84) coordinate system. This target's longitude and latitude position is the core target positioning output of the entire system. In engineering implementation, the target's longitude and latitude position serves as an output coordinate for: Annotation: Displayed on a custom-developed map model as a display for human-computer interaction; Task scheduling: Replanning subsequent drone routes and focusing tracking; Alert push: Generating alerts and communicating the coordinates to relevant personnel.
[0125] Step S214: splice the target longitude and latitude positions of the same target area at different timestamps to obtain a target trajectory sequence corresponding to each target area, and perform optimization processing on any target trajectory sequence. The optimization processing specifically includes:
[0126] (a) In the specific implementation, an observation model is constructed based on the target latitude and longitude positions contained in the target trajectory sequence. The observation model is used to optimize the target trajectory sequence through Kalman filtering to smooth the target trajectory sequence. In actual applications, due to factors such as sensor jitter, observation error, and attitude disturbance, the target trajectory sequence fluctuates. In order to suppress the jitter and positioning error in the target trajectory sequence, the target trajectory sequence is smoothed using Kalman filtering. Specifically:
[0127] State vector definition: ;in, Represents the state vector at timestamp The value of Waiting for the target state, Represents the horizontal position coordinate of the target area on the ground plane (ENU coordinate system). Since the ground target focused on in the embodiment of the present invention is usually assumed to be at the ground level, it is not necessary to track its position change in the vertical height direction z. Represents speed.
[0128] State transition equation: ;in, Represents the state vector at timestamp The value of is the state transition matrix, is the process noise.
[0129] Observation model: ;in, Represents the timestamp The observation value of the target state, 、 are process noise and observation noise respectively, both are zero-mean Gaussian noise, is the observation matrix.
[0130] Filtering process: includes two steps: prediction and update, refer to the linear Kalman filter process.
[0131] The embodiment of the present invention can improve the stability of the target trajectory sequence and reduce the probability of false alarms through Kalman filtering processing.
[0132] (b) For any timestamp in the target trajectory sequence, if the observation weight value corresponding to the timestamp is lower than the preset weight threshold, the target latitude and longitude position at the timestamp is removed from the target trajectory sequence.
[0133] (c) For any timestamp in the target trajectory sequence , determine the timestamp The target latitude and longitude position under With the previous timestamp The target latitude and longitude position under The spatial variation distance between , in spatially varying distances If the distance is greater than the preset threshold, the timestamp The target latitude and longitude position under Remove from the target trajectory sequence. Here, the timestamp is defined as The purpose is to compare with the timestamps involved in the above positioning solution process Make a distinction, timestamp Can be timestamped Same or different, 、 It is the abbreviation of the target's longitude and latitude position. In the specific implementation, if the target's longitude and latitude position deviation in consecutive frames exceeds the set value, it is judged as misidentification or occlusion and is removed or re-identified. Position deviation is the change in the spatial distance between the longitude and latitude positions of the same target in consecutive frames. ,Right now:
[0134] ;
[0135] Set a reasonable threshold here Defined as a preset physical reasonable offset (also known as a preset distance threshold), it is the maximum position deviation allowed by the drone's flight speed and frame rate. > When , the target latitude and longitude are judged to be misidentified or blocked and should be removed or re-identified.
[0136] (d) For any timestamp in the target trajectory sequence, a regional consistency check is performed on the target pixel position and the target latitude and longitude positions at that timestamp. If the regional consistency check fails, the target latitude and longitude positions at that timestamp are removed from the target trajectory sequence. In practice, the target pixel positions after fusion of the visible light image and the infrared image are subjected to geometric ray solving, multi-frame fusion, and least squares intersection to obtain the target latitude and longitude coordinates for each frame. This is then checked for regional consistency. The goal is to eliminate isolated pixels, small noise clumps, or pseudo targets with unreasonable geometric shapes. In actual engineering implementations, the detection strategies used include: Connected Domain Analysis: Determines whether the target is an isolated pixel or a small noise clump. The area threshold (number of pixels) must be greater than 10 pixels; Shape Rationality: Determines whether the target is an elongated, fragmented, or other shape that does not conform to the natural shape of the fire / smoke; Neighborhood Consistency: Checks whether there are related targets nearby and whether the target is separated from the overall fire / smoke area.
[0137] In summary, the multi-target fire location method based on multi-spectral dynamic fusion provided by the embodiment of the present invention includes at least the following key technical points: the embodiment of the present invention introduces a multi-factor joint calculation confidence scoring mechanism in the fire source identification and smoke identification processes, respectively, to achieve dynamic credibility evaluation of the identification results, and performs weighted fusion of the fire source mask area and the smoke mask area based on the confidence, forming a joint identification system oriented towards the task goal, and finally combines the target pixel position with pose solution, time synchronization, ray reconstruction, and least squares positioning to achieve precise positioning of the target area. The embodiment of the present invention can achieve the following effects through the integrated process of weighted coordinate fusion based on confidence, least squares intersection optimization, and multi-frame positioning and tracking under dynamic routes: high-precision identification and positioning of smoke and fire sources, significantly reducing the false detection and missed detection rate; adapting to dynamic route scenarios, without the need for hovering or fixed-point shooting, and significantly improving efficiency; supporting multi-target parallel processing to meet the needs of large-scale fire inspections; open architecture, which can be expanded to various application scenarios such as forest inspections and power equipment monitoring.
[0138] Based on the above embodiments, the present invention provides a multi-target fire location device based on multi-spectral dynamic fusion. Figure 3 The structure diagram of a multi-target fire location device based on multi-spectral dynamic fusion is shown. The device mainly includes the following parts:
[0139] An image acquisition module 302 is configured to acquire a multispectral image sequence collected by the UAV during an inspection mission. The multispectral image sequence includes multiple image frame groups, each of which includes a visible light image and an infrared image at the same timestamp.
[0140] Identification and fusion module 304 is configured to perform smoke identification processing, fire source identification processing, and dynamic fusion processing based on confidence levels based on the visible light image and infrared image contained in the image frame group, to obtain target pixel positions identified as target areas in the image frame group. The target area is the area where the smoke identification results and the fire source identification results are located together. There is at least one target area.
[0141] The position inversion module 306 is used to perform ray inversion based on the target pixel position identified as the target area in the image frame group, using the drone's posture data as positioning compensation to determine the target latitude and longitude position of the target area identified as the target area in the image frame group.
[0142] The multi-target fire positioning device based on multi-spectral dynamic fusion provided by the embodiment of the present invention performs smoke recognition, fire source recognition and dynamic fusion of confidence-based results through visible light images and infrared images to obtain the target pixel position of at least one target area identified as the target area. On this basis, ray inversion is performed using the posture data of the drone as positioning compensation to achieve precise positioning of multiple target areas and obtain the target longitude and latitude positions. The present invention can not only realize the identification and positioning of multiple target areas during the drone's cruise mission, which helps to improve the positioning efficiency of the target area, but also significantly improve the positioning accuracy of the target area.
[0143] In one embodiment, the identification and fusion module 304 is specifically configured to:
[0144] Performing smoke recognition processing on the visible light image at the current timestamp to obtain a smoke mask region, determining the smoke area and mask edge sharpness corresponding to the smoke mask region, and using this to obtain a smoke confidence level corresponding to the smoke mask region;
[0145] Performing fire source identification processing on the infrared image at the current timestamp to obtain a fire source mask area, determining the thermal feature change amplitude and mask stability corresponding to the fire source mask area, and obtaining the fire source confidence corresponding to the fire source mask area;
[0146] Based on the smoke confidence and the fire source confidence, the pixel position of the specified point in the smoke mask area and the pixel position of the specified point in the fire source mask area are fused to obtain the target pixel position identified as the target area at the current timestamp.
[0147] In one embodiment, the identification and fusion module 304 is specifically configured to:
[0148] Extract the highest and lowest temperature values within the fire source mask area at the current timestamp, and determine the thermal feature change amplitude corresponding to the fire source mask area of the current frame in combination with the standard temperature difference reference value;
[0149] and, determining the mask stability corresponding to the fire source mask area at the current timestamp based on the center of gravity change distance and mask area change ratio between the fire source mask area at the current timestamp and the fire source mask area at the previous timestamp;
[0150] The thermal feature variation amplitude and mask stability are weightedly fused to obtain the fire source confidence corresponding to the fire source mask area at the current timestamp.
[0151] In one embodiment, the position inversion module 306 is specifically configured to:
[0152] Based on the target pixel position identified as the target area at the current timestamp, and the motion parameters and gimbal angle contained in the drone's posture data, an observation ray vector pointing from the drone to the target area at the current timestamp is constructed;
[0153] Determine the observation weight values corresponding to multiple target timestamps, perform ray inversion on the current timestamp based on the observation ray vector and observation weight value at the target timestamp, and obtain the target latitude and longitude positions identified as the target area at the current timestamp.
[0154] In one embodiment, the position inversion module 306 is specifically configured to:
[0155] Based on the UAV's motion parameters, the UAV's position at the current timestamp is calculated;
[0156] According to the gimbal angle of the drone and the target pixel position identified as the target area at the current timestamp, the ray direction vector corresponding to the target pixel position at the current timestamp is determined;
[0157] Based on the ray direction vectors corresponding to the drone position and the target pixel position, the observation ray vector pointing from the drone to the target area at the current timestamp is constructed.
[0158] In one embodiment, the position inversion module 306 is specifically configured to:
[0159] For the target timestamp, perform the following operations:
[0160] Based on the corresponding values of the image clarity item, segmentation confidence item, attitude stability item and triangulation baseline perspective rationality item at the target timestamp, the observation weight value corresponding to the target timestamp is determined;
[0161] Among them, the image clarity item is used to describe the influence of the variance of the image recognition results of visible light images and infrared images on the observation weight value; the segmentation confidence item is used to describe the influence of the smoke confidence and fire source confidence on the observation weight value; the attitude stability item is used to describe the influence of the degree of change of the attitude angular velocity of the UAV on the observation weight value; the reasonable triangulation baseline perspective item is used to describe the influence of the angle between the UAV's perspective and the horizontal line and the length of the UAV's triangulation baseline on the observation weight value.
[0162] In one embodiment, the position inversion module 306 is specifically configured to:
[0163] By using a pre-built objective function, based on the observation ray vector and the observation weight value at the target timestamp, the target ground position identified as the target area at the current timestamp is determined so that the distance between the target ground position and the vertical position of the observation ray vector at the target timestamp is the shortest;
[0164] The target ground position is converted into coordinates to obtain the target latitude and longitude position identified as the target area at the current timestamp.
[0165] In one embodiment, a position stitching and trajectory optimization module is further included, which is used to:
[0166] The target longitude and latitude positions of the same target area at different timestamps are spliced together to obtain the target trajectory sequence corresponding to each target area. For any target trajectory sequence, perform at least one of the following optimization processes:
[0167] An observation model is constructed based on the target latitude and longitude positions contained in the target trajectory sequence, and the target trajectory sequence is optimized by Kalman filtering using the observation model to smooth the target trajectory sequence;
[0168] For any timestamp in the target trajectory sequence, if the observation weight value corresponding to the timestamp is lower than the preset weight threshold, the target latitude and longitude position at the timestamp is removed from the target trajectory sequence;
[0169] For any timestamp in the target trajectory sequence, determine the spatial variation distance between the target latitude and longitude position at that timestamp and the target latitude and longitude position at the previous timestamp. If the spatial variation distance is greater than a preset distance threshold, remove the target latitude and longitude position at that timestamp from the target trajectory sequence.
[0170] For any timestamp in the target trajectory sequence, a regional consistency check is performed on the target pixel position and the target latitude and longitude position at that timestamp. If the regional consistency check fails, the target latitude and longitude position at that timestamp is removed from the target trajectory sequence.
[0171] The device provided in the embodiment of the present invention has the same implementation principle and technical effects as those in the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the device embodiment, reference can be made to the corresponding content in the aforementioned method embodiment.
[0172] An embodiment of the present invention provides an electronic device. Specifically, the electronic device includes a processor and a storage device. The storage device stores a computer program, and when the computer program is executed by the processor, it executes the method described in any one of the above-mentioned embodiments.
[0173] Figure 4 This is a structural diagram of an electronic device provided in an embodiment of the present invention. The electronic device 100 includes: a processor 40, a memory 41, a bus 42 and a communication interface 43. The processor 40, the communication interface 43 and the memory 41 are connected via the bus 42; the processor 40 is used to execute an executable module stored in the memory 41, such as a computer program.
[0174] Memory 41 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. Communication between the system network element and at least one other network element is achieved through at least one communication interface 43 (which may be wired or wireless), and may utilize the Internet, a wide area network, a local area network, a metropolitan area network, or the like.
[0175] The bus 42 may be an ISA bus, a PCI bus, or an EISA bus. The bus may be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, Figure 4 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0176] Among them, the memory 41 is used to store programs, and the processor 40 executes the program after receiving the execution instruction. The method executed by the device for flow process definition disclosed in any embodiment of the above-mentioned embodiment of the present invention can be applied to the processor 40 or implemented by the processor 40.
[0177] Processor 40 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method may be completed by hardware integrated logic circuits or software instructions in processor 40. The above processor 40 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processing unit (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It may implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present invention may be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 41 , and the processor 40 reads the information in the memory 41 and completes the steps of the above method in combination with its hardware.
[0178] The computer program product of the readable storage medium provided in the embodiment of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method described in the previous method embodiment. The specific implementation can be referred to the previous method embodiment and will not be repeated here.
[0179] If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0180] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A multi-target fire location method based on multi-spectral dynamic fusion, characterized in that: include: Acquire a multispectral image sequence collected by the UAV during an inspection mission, wherein the multispectral image sequence includes a plurality of image frame groups, and the image frame groups include visible light images and infrared images at the same timestamp; Performing smoke recognition processing, fire source recognition processing, and dynamic fusion processing of the results based on confidence levels on the visible light image and the infrared image included in the image frame group to obtain a target pixel position identified as a target area in the image frame group, wherein the target area is an area where the results of the smoke recognition processing and the results of the fire source recognition processing are both located, and there is at least one target area. The confidence levels include a smoke confidence level and a fire source confidence level, wherein the smoke confidence level is determined based on the smoke area and mask edge sharpness corresponding to a smoke mask area, and the fire source confidence level is determined based on a thermal feature variation amplitude and mask stability corresponding to the fire source mask area. Based on the target pixel position identified as the target area in the image frame group, ray inversion is performed using the posture data of the drone as positioning compensation to determine the target latitude and longitude positions of the target area identified as the target area in the image frame group.
2. The multi-target fire location method based on multi-spectral dynamic fusion according to claim 1 is characterized in that: Performing smoke recognition processing, fire source recognition processing, and confidence-based dynamic fusion processing on the visible light image and the infrared image included in the image frame group to obtain a target pixel position identified as a target area in the image frame group includes: Performing smoke recognition processing on the visible light image at the current timestamp to obtain a smoke mask region, and determining a smoke area and mask edge sharpness corresponding to the smoke mask region to obtain a smoke confidence corresponding to the smoke mask region; Performing fire source identification processing on the infrared image at the current timestamp to obtain a fire source mask area, and determining a thermal feature variation amplitude and mask stability corresponding to the fire source mask area to obtain a fire source confidence level corresponding to the fire source mask area; Based on the smoke confidence and the fire source confidence, the pixel position of the designated point in the smoke mask area and the pixel position of the designated point in the fire source mask area are fused to obtain the target pixel position identified as the target area at the current timestamp.
3. The multi-target fire location method based on multi-spectral dynamic fusion according to claim 2 is characterized in that: Determining the thermal signature variation amplitude and mask stability corresponding to the fire source mask area to obtain the fire source confidence corresponding to the fire source mask area includes: Extracting the highest temperature value and the lowest temperature value in the fire source mask area at the current timestamp, and determining the thermal feature change amplitude corresponding to the fire source mask area of the current frame in combination with a standard temperature difference reference value; and determining the mask stability corresponding to the fire source mask area at the current timestamp based on a center of gravity change distance and a mask area change ratio between the fire source mask area at the current timestamp and the fire source mask area at the previous timestamp; The thermal feature variation amplitude and the mask stability are weightedly fused to obtain the fire source confidence corresponding to the fire source mask area at the current timestamp.
4. The multi-target fire location method based on multi-spectral dynamic fusion according to claim 1 is characterized in that: Based on the target pixel position identified as the target area in the image frame group, ray inversion is performed using the posture data of the drone as positioning compensation to determine the target latitude and longitude position of the target area identified as the target area in the image frame group, including: Based on the target pixel position identified as the target area at the current timestamp, and the motion parameters and gimbal angle included in the posture data of the drone, construct an observation ray vector directed from the drone to the target area at the current timestamp; Determine observation weight values corresponding to multiple target timestamps, perform ray inversion on the current timestamp based on the observation ray vector and the observation weight value at the target timestamp, and obtain the target latitude and longitude position identified as the target area at the current timestamp.
5. The multi-target fire location method based on multi-spectral dynamic fusion according to claim 4 is characterized in that: Based on the target pixel position identified as the target area at the current timestamp, and the motion parameters and the gimbal angle included in the posture data of the drone, an observation ray vector directed from the drone to the target area at the current timestamp is constructed, including: Calculating the drone position of the drone at a current timestamp based on the motion parameters of the drone; Determine, based on the gimbal angle of the drone and the target pixel position identified as the target area at the current timestamp, a ray direction vector corresponding to the target pixel position at the current timestamp; Based on the ray direction vector corresponding to the drone position and the target pixel position, an observation ray vector directed from the drone to the target area at the current timestamp is constructed.
6. The multi-target fire location method based on multi-spectral dynamic fusion according to claim 4 is characterized in that: Determine the observation weight values corresponding to multiple target timestamps, including: For any target timestamp, perform the following operations: Determining an observation weight value corresponding to the target timestamp based on the corresponding values of the image clarity item, the segmentation confidence item, the posture stability item, and the triangulation baseline perspective rationality item at the target timestamp; Among them, the image clarity item is used to describe the influence of the variance of the image recognition results of the visible light image and the infrared image on the observation weight value; the segmentation confidence item is used to describe the influence of the smoke confidence and the fire source confidence on the observation weight value; the attitude stability item is used to describe the influence of the degree of change of the attitude angular velocity of the UAV on the observation weight value; the reasonable triangulation baseline perspective item is used to describe the influence of the angle between the perspective of the UAV and the horizontal line and the length of the triangulation baseline of the UAV on the observation weight value.
7. The multi-target fire location method based on multi-spectral dynamic fusion according to claim 4 is characterized in that: Performing ray inversion based on the observation ray vector and the observation weight value at the target timestamp to obtain the target latitude and longitude position identified as the target area at the current timestamp includes: determining, by a pre-constructed objective function, a target ground position identified as the target area at the current timestamp based on the observation ray vector at the target timestamp and the observation weight value, such that a vertical distance between the target ground position and the observation ray vector at the target timestamp is the shortest; The target ground position is coordinate-converted to obtain the target longitude and latitude position identified as the target area at the current timestamp.
8. The multi-target fire location method based on multi-spectral dynamic fusion according to claim 1 is characterized in that: The method further comprises: The target longitude and latitude positions of the same target area at different timestamps are spliced to obtain a target trajectory sequence corresponding to each target area, and at least one of the following optimization processes is performed on any target trajectory sequence: constructing an observation model based on the target latitude and longitude positions contained in the target trajectory sequence, and performing Kalman filter optimization on the target trajectory sequence using the observation model to smooth the target trajectory sequence; For any timestamp in the target trajectory sequence, if the observation weight value corresponding to the timestamp is lower than the preset weight threshold, the target latitude and longitude position at the timestamp is removed from the target trajectory sequence; For any timestamp in the target trajectory sequence, determining a spatial variation distance between the target longitude and latitude position at the timestamp and the target longitude and latitude position at the previous timestamp, and removing the target longitude and latitude position at the timestamp from the target trajectory sequence if the spatial variation distance is greater than a preset distance threshold; For any timestamp in the target trajectory sequence, a regional consistency check is performed on the target pixel position and the target latitude and longitude position at the timestamp. If the regional consistency check fails, the target latitude and longitude position at the timestamp is removed from the target trajectory sequence.
9. A multi-target fire location device based on multi-spectral dynamic fusion, characterized in that: include: An image acquisition module is used to acquire a multispectral image sequence collected by the UAV during the inspection mission, wherein the multispectral image sequence includes multiple image frame groups, and the image frame groups include visible light images and infrared images at the same time stamp; an identification and fusion module, configured to perform smoke identification processing, fire source identification processing, and dynamic fusion processing of the results based on confidence, based on the visible light image and the infrared image included in the image frame group, to obtain a target pixel position identified as a target area in the image frame group, wherein the target area is an area where the results of the smoke identification processing and the results of the fire source identification processing are both located, and there is at least one target area. The confidence includes a smoke confidence factor and a fire source confidence factor, wherein the smoke confidence factor is determined based on the smoke area corresponding to the smoke mask area and the sharpness of the mask edge, and the fire source confidence factor is determined based on the thermal feature variation amplitude and mask stability corresponding to the fire source mask area; A position inversion module is used to perform ray inversion based on the target pixel position identified as the target area in the image frame group, using the posture data of the drone as positioning compensation, to determine the target latitude and longitude position identified as the target area in the image frame group.
10. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the method according to any one of claims 1 to 8.