Multi-spectral fire source location accurate detection method under complex fire scene

CN122591659APending Publication Date: 2026-08-18GUANGDONG POLYTECHNIC OF IND & COMMERCE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610802280.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-04
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

然而,当前的探测机制往往陷入追求极致物理速度的误区,单纯依靠提升云台电机的运转速度来缩短切换时间

Benefits of technology

本发明公开了一种复杂火灾场景下多光谱火源位置精准探测方法,针对云台高速转动导致成像设备机械振动引发图像模糊从而影响火源定位精度的核心问题,创新性地采用长短期记忆网络预测不同转速与设备质量分布组合下的残余振动幅度,在云台到达目标位置前预判振动风险并自适应调整转速,有效抑制了机械振动对成像质量的干扰。本发明进一步通过相邻帧边缘轮廓吻合度分析识别画面晃动程度,对超阈值帧实施去模糊处理获得稳定图像组,结合支持向量机识别火源燃烧阶段并动态确定观测窗口时长,确保采集足够数据帧后通过火焰质心坐标反投影实现火源三维空间精准定位,显著提升了复杂火灾场景下火源位置探测的准确性与可靠性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122591659A_ABST
    Figure CN122591659A_ABST
Patent Text Reader

Abstract

This application provides a method for accurate multispectral fire source location detection in complex fire scenarios, comprising: scanning the entire monitoring area using an infrared imaging component mounted on a gimbal, identifying areas with abnormal temperatures and locating the range of each target fire source, and setting an initial gimbal rotation speed accordingly; inputting the initial gimbal rotation speed and the mass distribution data of the detection device into a pre-established long short-term memory network to obtain a predicted result of the mechanical vibration amplitude when the detection device reaches the target fire source location; when the predicted mechanical vibration amplitude exceeds a preset vibration threshold, reducing the initial gimbal rotation speed according to a preset attenuation ratio to obtain the target gimbal rotation speed; segmenting the flame area from a stable image group, extracting the flame combustion area and temperature distribution, identifying the current fire source combustion stage using a support vector machine, calculating the minimum number of data frames required for reliable updates of the combustion stage, and determining the observation window duration for the current target fire source location.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information technology, and in particular to a method for accurate detection of multispectral fire source locations in complex fire scenarios. Background Technology

[0002] Precise multispectral fire source location detection in complex fire scenarios is a crucial aspect of modern fire rescue and disaster control. Faced with emergencies where multiple fire sources spread simultaneously, detection equipment must quickly switch between different sources to obtain a comprehensive fire situation. However, current detection mechanisms often fall into the trap of pursuing extreme physical speed, simply relying on increasing the operating speed of the pan-tilt motor to shorten the switching time. When the pan-tilt rotates at extremely high speed and suddenly stops at the target fire source location, the enormous mechanical inertia inevitably generates strong residual mechanical vibrations, causing severe shaking of the equipment's lens. Although it can quickly align with the target, it cannot obtain high-quality fire images. Waiting for the mechanical vibrations to completely subside requires additional time, which in turn encroaches on the time originally intended for capturing the fire source. For example, when multiple fires occur in a large storage area or chemical production area, the fire sources may be distributed in areas with vastly different spatial locations, such as raw material storage areas, production and processing areas, and finished product loading and unloading areas. The detection equipment must sequentially poll the three relatively far-away fire source locations within a specified global update cycle. The faster the pan-tilt unit rotates, the more violent the braking vibration after reaching each target location, and the longer the time consumed to wait for the image to stabilize. As a result, the effective dwell window is greatly compressed, and the flame temperature field in the infrared band and the combustion contour in the visible light band cannot be completely collected with sufficient clear frames. This makes it impossible for the system to accurately judge the current combustion stage and spread of each fire source. Summary of the Invention

[0003] This invention provides a method for accurate detection of multispectral fire source locations in complex fire scenarios, mainly including: The infrared imaging component mounted on the PTZ camera scans the entire monitoring area, identifies areas with abnormal temperatures, and locates the range of each target fire source, thereby setting the initial PTZ camera rotation speed. The initial gimbal rotation speed and the mass distribution data of the detection device are input into a pre-established long short-term memory network to obtain the predicted mechanical vibration amplitude when the detection device reaches the target fire source location. When the predicted mechanical vibration amplitude exceeds the preset vibration threshold, the initial gimbal rotation speed is downgraded according to the preset attenuation ratio to obtain the target gimbal rotation speed. After the target gimbal rotation speed drives the detection device to reach the target fire source location, it acquires image frame groups, analyzes the edge contours of adjacent frames, determines the degree of image shaking by identifying the degree of matching of the edge contours of adjacent frames, and performs deblurring on frames whose image shaking degree exceeds the preset shaking threshold to obtain a stable image group. The flame region is segmented from the stable image group, the flame combustion area and temperature distribution are extracted, the current combustion stage of the fire source is identified by support vector machine, the minimum number of data frames required to update the combustion stage is calculated, and the observation window duration of the current target fire source location is determined. Stable image frames are continuously acquired within the observation window duration. After accumulating to the minimum number of data frames, the centroid pixel coordinates of the flame outline in each frame are extracted. The centroid pixel coordinates are then mapped to the three-dimensional coordinate system of the monitoring area through coordinate back projection to obtain the three-dimensional coordinates of the current target fire source.

[0004] Furthermore, the process of using the infrared imaging component mounted on the pan-tilt unit to perform a full-range scan of the monitored area, identify areas with abnormal temperatures, and locate the range of each target fire source, thereby setting the initial pan-tilt unit rotation speed, includes: The monitoring area is scanned line by line by the infrared imaging component mounted on the pan-tilt unit to obtain the radiation intensity of each pixel. Continuous pixels with radiation intensity higher than a preset abnormal threshold are clustered to obtain the temperature abnormal area. The temperature anomaly area is extended outward along the direction of radiation intensity from high to low until the radiation intensity drops back to the background radiation intensity, thus defining the outline boundary of the range of each target fire source. Based on the outline boundary of each target fire source range, extract the centroid azimuth and elevation angle of each fire source, and determine the polling order of the gimbal to arrive at each target fire source in sequence according to the azimuth angle. Based on the distance between two adjacent target fire sources in the polling order, a higher speed gear is matched for polling segments with a distance greater than a preset distance threshold, and a lower speed gear is matched for polling segments with a distance less than the preset distance threshold, so as to obtain the initial gimbal speed of each polling segment.

[0005] Furthermore, the step of extracting the centroid azimuth and pitch angle of each fire source and determining the polling order of the gimbal to arrive at each target fire source in sequence according to the azimuth angle includes: calculating the azimuth distance span between adjacent fire sources for each pair of target fire sources based on the spatial distribution of each target fire source, arranging each target fire source in ascending order of azimuth angle, and determining the polling order of the gimbal to arrive at each target fire source in sequence.

[0006] Furthermore, the step of inputting the initial gimbal rotation speed and the mass distribution data of the detection device into a pre-established long short-term memory network to obtain the predicted mechanical vibration amplitude when the detection device reaches the target fire source location includes: Based on the mass distribution data, the weight distribution of each segment of the gimbal robotic arm along the arm length direction and the rotational inertia parameters of the imaging component at the installation position are extracted. The initial gimbal rotation speed, weight distribution and rotational inertia parameters are arranged according to the continuous rotation time before the gimbal reaches the target fire source position to obtain a time sequence reflecting the rotation process. The residual flutter waveforms of the detection equipment after reaching the target point with different initial gimbal rotation speeds and different mass distributions during its historical operation are collected. The vibration amplitude decay curve from the moment of pause is extracted. The time series is used as input and the peak amplitude of the vibration amplitude decay curve is used as output to label training samples and iteratively update the weights of the internal memory units of the long short-term memory network. The time sequence is fed into the pre-established long short-term memory network step by step, and the peak amplitude of residual flutter is recursively output based on the rotational inertia parameter and speed change at the pause time to determine the mechanical vibration amplitude prediction result.

[0007] Furthermore, when the predicted mechanical vibration amplitude exceeds a preset vibration threshold, the initial gimbal rotation speed is downgraded according to a preset attenuation ratio to obtain the target gimbal rotation speed. This includes: comparing the residual flutter peak amplitudes in the predicted mechanical vibration amplitude with the preset vibration threshold one by one; marking the polling segment where the target fire source location exceeds the preset vibration threshold as the polling segment to be downgraded; reducing the initial gimbal rotation speed of the polling segment step by step according to the preset attenuation ratio; and re-outputting the residual flutter peak amplitude by the long short-term memory network for each reduction until it falls below the preset vibration threshold to obtain the target gimbal rotation speed.

[0008] Furthermore, after the target gimbal rotation speed drives the detection device to reach the target fire source location, it acquires a group of image frames, analyzes the edge contours of adjacent frames, determines the degree of image shake by identifying the degree of matching of the edge contours of adjacent frames, and performs deblurring on frames whose image shake exceeds a preset shake threshold to obtain a stable image group, including: Activate the infrared and visible light dual-band imaging component, continuously record fire source images at fixed acquisition times, and extract the edge contours at the boundary between the flame and the background for each frame. The pixel offset is calculated for the corresponding point of the same flame edge in two adjacent frames to obtain the contour displacement of the two adjacent frames. The degree of matching of the edge contour of the adjacent frames is determined based on the magnitude of the contour displacement, and the degree of image shaking is determined based on the degree of matching. For frames whose image shake exceeds a preset shake threshold, Wiener filtering is used to construct a point spread function based on the frame contour displacement direction and displacement magnitude to deblur each pixel. The clear frames after deblurring are then merged with the frames that do not exceed the preset shake threshold according to the acquisition time to obtain a stable image group.

[0009] Furthermore, the step of segmenting the flame region from the stable image group, extracting the flame combustion area and temperature distribution, identifying the current fire source combustion stage through a support vector machine, calculating the minimum number of data frames required for updating the combustion stage, and determining the observation window duration for the current target fire source location includes: extracting closed contours to delineate the flame region for visible light band frames, calculating the flame combustion area by counting the total number of pixels, and calculating the temperature distribution by reading the radiation intensity of each pixel within the flame region for infrared band frames; concatenating the statistics of the flame combustion area and temperature distribution to form a feature vector; inputting the feature vector into the support vector machine, determining the combustion stage based on the region corresponding to each combustion stage; determining the minimum number of data frames by converging the fluctuation amplitude of adjacent discrimination results of combustion stages to within a preset fluctuation threshold, and dividing this number by the number of frames acquired per second by the imaging component to determine the observation window duration.

[0010] Furthermore, the step of inputting the feature vector into the support vector machine and determining the combustion stage based on the region corresponding to each combustion stage includes: the support vector machine using the feature vectors of historical fire sources in the initial combustion, development and intense combustion stages as training samples to obtain the support vectors and classification hyperplanes that divide each combustion stage, and determining the combustion stage of the current fire source based on the region where the feature vector falls.

[0011] Furthermore, within the observation window duration, stable image frames are continuously acquired. After accumulating to the minimum number of data frames, the centroid pixel coordinates of the flame outline in each frame are extracted. The centroid pixel coordinates are then mapped to the three-dimensional coordinate system of the monitoring area through coordinate back projection to obtain the three-dimensional coordinates of the current target fire source. This includes: accumulating and counting stable image frames until the minimum number of data frames is reached to obtain a stable frame group; extracting the flame outline for each frame, calculating the geometric center of the surrounding pixels to obtain the centroid pixel coordinates, and taking the average value according to the frame order to obtain the stable centroid pixel coordinates; acquiring camera parameters and the flame area depth distance read by the infrared band, and using coordinate back projection to map the stable centroid pixel coordinates to the three-dimensional coordinate system of the monitoring area to obtain the three-dimensional coordinates.

[0012] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects: This invention discloses a method for accurate multispectral fire source location detection in complex fire scenarios. Addressing the core issue of image blurring caused by mechanical vibration of the imaging equipment due to high-speed gimbal rotation, which affects fire source location accuracy, this invention innovatively employs a long short-term memory network to predict the residual vibration amplitude under different combinations of rotation speed and equipment mass distribution. Vibration risk is predicted before the gimbal reaches the target position, and the rotation speed is adaptively adjusted, effectively suppressing the interference of mechanical vibration on image quality. Furthermore, this invention identifies the degree of image shake through adjacent frame edge contour matching analysis, performs deblurring on frames exceeding a threshold to obtain stable image sets, and combines support vector machines to identify the fire source combustion stage and dynamically determine the observation window duration. This ensures that after acquiring sufficient data frames, accurate three-dimensional spatial positioning of the fire source is achieved through back-projection of the flame centroid coordinates, significantly improving the accuracy and reliability of fire source location detection in complex fire scenarios. Attached Figure Description

[0013] Figure 1 This is a flowchart of a method for accurate detection of multispectral fire source location in complex fire scenarios according to the present invention.

[0014] Figure 2 This is a schematic diagram of a method for accurate detection of multispectral fire source location in complex fire scenarios according to the present invention.

[0015] Figure 3 This is another schematic diagram of a method for accurate detection of multispectral fire source location in complex fire scenarios according to the present invention. Detailed Implementation

[0016] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] like Figures 1-3 This embodiment of a method for accurate detection of multispectral fire source location in complex fire scenarios may specifically include: S101. The monitoring area is scanned in full range by the infrared imaging component mounted on the pan-tilt unit, the abnormal temperature area is identified and the range of each target fire source is located, and the initial pan-tilt unit speed is set accordingly.

[0018] The monitoring area is scanned line by line by the infrared imaging component mounted on the pan-tilt unit to obtain the radiation intensity of each pixel. Pixels with radiation intensities exceeding a preset anomaly threshold are clustered to identify temperature anomaly regions. These temperature anomaly regions are then extended outwards along a direction of decreasing radiation intensity until the radiation intensity returns to the background intensity, defining the contour boundaries of each target fire source's range. Based on these contour boundaries, the centroid azimuth and elevation angles of each fire source within the monitoring area are extracted to obtain the spatial distribution of each target fire source. The azimuth spacing between adjacent fire sources is calculated pairwise for each fire source in this spatial distribution. The target fire sources are arranged in ascending order of azimuth angle to determine the polling order in which the pan-tilt unit reaches each target fire source. Based on the spacing between adjacent fire sources in the polling order, if a single spacing span is greater than a preset span threshold, a higher rotation speed is matched; if a single spacing span is less than the preset span threshold, a lower rotation speed is matched, resulting in the initial pan-tilt unit rotation speed for each polling segment.

[0019] A multi-fire source detection device for complex fire scenarios uses a pan-tilt unit (PTZ) as its main support. An infrared imaging component is fixedly mounted at the front end of the PTG's robotic arm. This component collects radiation signals in the 8-14 micrometer long-wave infrared band, which corresponds to the peak energy range of thermal radiation from the surface of a burning object. Flames and high-temperature plumes exhibit significantly higher radiation intensity than the environmental background in this band. After power-on, the infrared imaging component moves with the PTG in a two-dimensional angular space defined by the horizontal and vertical directions, providing full coverage of the monitored area. In one embodiment, the full-range scanning of the monitored area employs a line-by-line reciprocating method. The PTG rotates uniformly from the starting azimuth angle to the ending azimuth angle at a fixed elevation angle. After completing one line of scanning, the elevation angle increases by one field of view angle, and then the next line is scanned in reverse, reciprocating until the entire monitored area is covered. During the scanning process, the infrared imaging component outputs thermal radiation image frames at fixed time intervals. Each pixel in each frame carries a radiation intensity reading, reflecting the surface temperature of the object in the corresponding spatial direction.

[0020] Specifically, the determination of temperature anomaly areas is based on a comparison between radiation intensity and a preset anomaly threshold. The preset anomaly threshold is taken as a several times the ambient background radiation intensity.

[0021] For example, the background radiation intensity of materials at room temperature in the raw material storage area corresponds to a surface temperature of approximately 30 degrees Celsius, while the surface temperature of the initial combustion source is generally higher than 200 degrees Celsius. Therefore, the abnormal threshold can be set to a radiation intensity reading corresponding to 120 degrees Celsius. For a single frame image, pixels with radiation intensities higher than this abnormal threshold are marked as abnormal pixels. Then, clustering is performed based on the adjacency relationship between pixels: if two abnormal pixels are adjacent in eight directions on the image, they are assigned to the same cluster. After merging them one by one, several unconnected continuous pixel clusters are obtained, and each cluster represents a temperature abnormality area.

[0022] It should be noted that in multi-point fires in storage areas or chemical production areas, independent temperature anomaly zones may simultaneously appear in raw material storage areas, production and processing areas, and finished product loading and unloading areas. Clustering naturally breaks down spatially separated fire sources into independent clusters, providing a basis for subsequent individual location. Further, the contour boundaries of these temperature anomaly zones are defined. Starting from the pixel with the highest radiation intensity within a cluster, adjacent pixels are detected outwards in a circle along the direction of decreasing radiation intensity. Pixels whose radiation intensity is still higher than the background radiation intensity are included in the fire source range until the radiation intensity of adjacent pixels drops back to the background radiation intensity. The line connecting these pixels that stops extending constitutes the contour boundary of the target fire source range.

[0023] Preferably, the contour boundary is recorded as a closed polygon, showing its pixel position in the image coordinate system. Based on the above contour boundary, the spatial distribution of each target fire source is extracted.

[0024] Specifically, the geometric centroid of each closed contour boundary is determined. The image pixel coordinates of the centroid are combined with the azimuth and elevation angles of the gimbal when the infrared imaging component acquires the frame. The azimuth and elevation angles of the fire source centroid relative to the device are then calculated. The set of azimuth and elevation angles of all fire sources constitutes the spatial distribution of each target fire source.

[0025] Understandably, the azimuth spacing span reflects the angle the gimbal needs to turn from one fire source to another. For the aforementioned spatial distribution, the azimuth angles of any two fire sources are subtracted and their absolute values ​​are taken; similarly, the pitch angles are subtracted, and the sum of these two values ​​yields the azimuth spacing span between the two fire sources. In another embodiment, the azimuth spacing span between adjacent fire sources is calculated pairwise for the aforementioned spatial distribution, and the azimuth angles of the i-th and (i+1)-th fire sources are denoted as α and β, respectively. i With a i+1 Then the azimuth spacing span d i =|a i+1 -a i First, the target fire sources are initially arranged in ascending order of azimuth angle. Then, the initial sequence is locally adjusted using nearest neighbor check. That is, if the sum of the spans of two adjacent segments is greater than the total span after the segment swap, the order of the two fire sources is swapped to determine the polling order of the pan-tilt unit to arrive at each target fire source in sequence, so as to ensure that the total scanning path is short.

[0026] In one embodiment, the azimuth angle difference between the fire source in the finished product loading and unloading area of ​​the chemical production zone and the fire source in the raw material stacking area can reach 130 degrees, while the azimuth angle difference between two fire sources on adjacent shelves in the production and processing area is only 15 degrees, resulting in a significant difference in the azimuth distance span. Subsequently, all target fire sources are sorted from smallest to largest azimuth angle. The gimbal then arrives at each fire source sequentially according to the sorting result. A polling segment is formed between two adjacent positions, and all polling segments are connected in series to form a complete polling sequence, avoiding repeated back-and-forth movements of the gimbal. The initial gimbal rotation speed for each polling segment is determined, and the initial gimbal rotation speed is matched to the corresponding speed level based on the comparison result between the distance span between two adjacent target fire sources in the polling segment and a preset span threshold.

[0027] For example, a preset span threshold of 60 degrees is used. If the span of a single segment in a polling segment is greater than this threshold, it indicates that the two fire sources are far apart, and a higher rotation speed is matched for that segment. If the span of a single segment is less than this threshold, it indicates that the two fire sources are close together, and a lower rotation speed is matched. After matching segment by segment, the initial gimbal rotation speed of the gimbal in each polling segment is obtained. This initial gimbal rotation speed is used as the reference for the rotation speed of the gimbal when polling each target fire source.

[0028] S102. Input the initial gimbal rotation speed and the mass distribution data of the detection equipment into a pre-established long short-term memory network to obtain the predicted mechanical vibration amplitude when the detection equipment reaches the target fire source location.

[0029] Using the gimbal rotation speed at the moment of initial rotation as a baseline, the instantaneous rotation speed at subsequent moments is recorded at 10-millisecond sampling intervals. This is aligned with the segmented weight distribution of the robotic arm and the rotational inertia parameters of the infrared and visible light imaging components at their installation positions, forming a time-series sequence spanning the entire rotation process until the moment of cessation. Residual flutter waveforms of the detection equipment after reaching the target point with different baseline rotation speeds and mass distributions during historical operations are collected. The vibration amplitude decay curve from the moment of cessation is extracted. Using the time-series sequence as input samples and the peak amplitude of the decay curve as labels, the rotation state is input to the Long Short-Term Memory (LSTM) network moment by moment, and the weights of the internal memory units are iteratively updated, completing the pre-training of the network. During the online inference phase, the Long Short-Term Memory (LSTM) network receives the current moment of inertia parameter J and the rotational speed increment Δw at each sampling moment. The cell state is updated jointly by the forget gate, input gate, and output gate, and the current hidden state is passed forward. When the time sequence is input to the pause moment, the network locks the hidden state vector at that moment and maps it to a scalar A through a fully connected layer and a linear activation function. A is the peak amplitude of the residual flutter, where J reflects the inertial energy storage level and Δw reflects the momentum mutation intensity. The two are cumulatively determined by the memory unit to determine the output magnitude of A, and finally the prediction result of the mechanical vibration amplitude when the detection device reaches the target fire source location is obtained.

[0030] In complex fire scenarios, when the detection equipment reaches the target fire source, the high-speed rotation of the pan-tilt unit causes residual vibration due to mechanical inertia during the moment of pause, resulting in camera shake. The severity of the residual vibration is closely related to the rotation state and the mass distribution of the equipment itself. Therefore, predicting the amplitude of mechanical vibration after reaching the target point based on the initial pan-tilt unit rotation speed and mass distribution data before the actual rotation of the pan-tilt unit constitutes the core of this solution.

[0031] Specifically, the mass distribution data refers to the spatial weight distribution and rotational inertia parameters of the gimbal robotic arm and its imaging components. The gimbal robotic arm is a slender structure, divided into several segments along its length from the root to the end of the axis of rotation, each segment having an independent weight reading. The infrared imaging component and the visible light imaging component are fixed at different mounting positions on the robotic arm, with their mass concentrated at the mounting point. Their distance from the axis of rotation determines the rotational inertia contributed by each component. The rotational inertia parameter reflects an object's ability to resist changes in its rotational state. The farther a component is from the axis of rotation and the greater its mass, the greater its rotational inertia, the stronger the accumulated inertial torque at rest, and the more severe the residual vibration. In one embodiment, when extracting the mass distribution data, the weight of each segment of the robotic arm is multiplied by the square of its distance from the axis of rotation, and then accumulated segment by segment. The rotational inertia components of the two imaging components at their respective mounting positions are then superimposed to obtain the overall rotational inertia parameter of the detection device around the axis of rotation. This rotational inertia parameter, together with the initial gimbal rotation speed, determines the inertial energy level at the moment of rest.

[0032] It should be noted that residual vibration does not occur all at once at the moment of pause, but rather is a process that gradually decays over time, and its peak value is affected by the continuous rotation process of the gimbal before it reaches the target fire source position. The rotational speed value at a single moment is insufficient to characterize this process. Therefore, the initial gimbal rotational speed, the weight distribution of each segment of the robotic arm, and the rotational inertia parameters of the two imaging components are arranged in the order of the continuous rotation moments before the gimbal reaches the target fire source position, forming a time sequence reflecting the rotation process. Each moment in the time sequence corresponds to a set of rotational state readings.

[0033] It is understandable that the time sequence exhibits interdependence between consecutive moments; the rotational state at an earlier moment influences the inertia accumulation at a later moment. This temporal dependency aligns perfectly with the processing characteristics of Long Short-Term Memory (LSTM) networks. LSTM is a recurrent neural network that processes sequential data. It contains memory units that control the retention and discarding of information through input gates, forget gates, and output gates. This allows the rotational state information from earlier moments to be transmitted along the temporal direction and accumulated up to the pause point, thereby capturing the cumulative contribution of the rotational process to residual tremors.

[0034] In one embodiment, the training process of the Long Short-Term Memory (LSTM) network relies on historical operational data of the detection device. After reaching the target point with different initial gimbal rotation speeds and mass distributions during historical polling, the vibration sensors on the detection device's lens or body record residual vibration waveforms after the pause. These residual vibration waveforms are curves showing the amplitude of vibration fluctuating over time. For this waveform, a decay curve showing the gradual decrease in vibration amplitude is extracted from the pause moment, and the peak amplitude of the decay curve is read. This peak amplitude represents the most severe shaking of the image after the pause. Each historical operational time sequence is used as the model input, and the corresponding peak amplitude of the vibration amplitude decay curve is used as the model output, forming labeled training samples. Furthermore, the model training employs a time-sequence input method. A time sequence of a training sample is fed into the Long Short-Term Memory (LSTM) network in chronological order. At each time step, the memory unit receives the current rotation state and updates its internal storage. At the pause point, a predicted peak value is output. The predicted peak value is compared with the actual peak value amplitude labeled in the sample. The difference between the two is propagated backward along the time direction. The weights of the input gate, forget gate, and output gate in the memory unit are iteratively adjusted. This process is repeated until the deviation between the predicted peak value and the actual peak value converges to below a preset deviation threshold, thus obtaining the pre-established LSM network.

[0035] For example, in a multi-fire-source polling scenario within a chemical production area, when the gimbal rotates at a relatively high initial rotational speed towards a fire source in a distant finished product loading / unloading area, the residual vibration peak value corresponding to this type of rotation is higher in historical samples due to the large rotational inertia of the visible light imaging component at the end of the robotic arm. During training, the model learns the mapping relationship between this type of rotational speed and mass distribution combination, resulting in a higher peak value upon reaching the target point. After training is complete, for the target fire source to be polled, its corresponding time series is fed into the pre-established Long Short-Term Memory network moment by moment. The network recursively outputs the peak amplitude of the residual vibration based on the rotational inertia parameters and rotational speed changes at pause points in the time series, thus determining the predicted mechanical vibration amplitude when the detection equipment reaches the target fire source location.

[0036] Preferably, the above prediction is performed on each target fire source in the raw material storage area, production and processing area and finished product loading and unloading area to obtain the predicted mechanical vibration amplitude corresponding to the location of each target fire source, which provides a basis for subsequent determination of whether to reduce the initial gimbal speed.

[0037] S103. When the predicted mechanical vibration amplitude exceeds the preset vibration threshold, the initial gimbal speed is downgraded according to the preset attenuation ratio to obtain the target gimbal speed.

[0038] The residual flutter peak amplitude corresponding to each target fire source location in the mechanical vibration amplitude prediction results is obtained. The residual flutter peak amplitude is compared one by one with a preset vibration threshold. If the residual flutter peak amplitude exceeds the preset vibration threshold, the polling segment containing the target fire source location is marked as a polling segment to be downgraded, thus obtaining an over-limit judgment result. For the polling segment to be downgraded in the over-limit judgment result, the current initial gimbal rotation speed of the polling segment is obtained. The initial gimbal rotation speed is gradually reduced according to a preset attenuation ratio. For each reduction, the residual flutter peak amplitude of the polling segment to be downgraded is re-output using the Long Short-Term Memory network until the residual flutter peak amplitude falls below the preset vibration threshold, thus obtaining the target gimbal rotation speed of the polling segment to be downgraded.

[0039] Specifically, for the polling segment to be downgraded in the over-limit judgment result, the initial gimbal speed v0 of the current polling segment is obtained. The initial gimbal speed is reduced step by step according to a fixed attenuation ratio of 10% per level. The speed after the k-th level reduction is denoted as vk = v0 * 0.9^k, where k is the number of reduction levels. After each level of reduction is completed, the updated speed and the feature vector corresponding to the polling segment are re-input into the Long Short-Term Memory network, and the residual jitter peak amplitude is re-output and compared with the preset vibration threshold of 0.8 mm / s. The maximum number of reduction levels is set to 5 levels. If the residual jitter peak amplitude falls below 0.8 mm / s for the first time at a certain level, the iteration is immediately terminated and the vk of the level reduction is taken as the target gimbal speed for the polling segment to be downgraded. If the target gimbal speed is still not met after the 5th level reduction, the target gimbal speed is forcibly locked at 50% of v0 as a fallback value, and an alarm flag is triggered to notify manual review, so as to avoid infinite reduction or inspection failure due to excessively low speed. For example, if the initial rotation speed of a polling segment to be downgraded is 60 degrees per second, and the predicted residual jitter peak amplitude still exceeds the threshold when the first level is reduced to 54 degrees per second, and the predicted value drops to 0.72 millimeters per second when the second level is reduced to 48.6 degrees per second, then the iteration is terminated and 48.6 degrees per second is determined as the target gimbal rotation speed for that polling segment. In multi-fire-source polling in complex fire scenarios, the mechanical vibration amplitude prediction result gives the residual jitter peak amplitude corresponding to the location of each target fire source. This peak amplitude characterizes the most severe shaking of the image after the gimbal reaches the target point and stops at the initial gimbal rotation speed. When the peak amplitude is too large, the lens image shakes violently, and neither the flame temperature field in the infrared band nor the combustion contour in the visible light band can be captured as a clear frame. Therefore, before driving the gimbal, it is determined whether to downgrade the initial gimbal rotation speed based on this peak amplitude. The preset vibration threshold is the upper limit of the residual jitter peak amplitude allowed for the lens image to still capture a clear frame.

[0040] For example, once the lens imaging resolution is determined, if the pixel displacement caused by shaking between two adjacent frames exceeds 3 pixels, obvious ghosting appears at the edge of the image. Based on this, the residual vibration amplitude corresponding to this displacement is set as a preset vibration threshold. The residual vibration peak amplitude corresponding to each target fire source location in the mechanical vibration amplitude prediction result is compared one by one with this preset vibration threshold. In one embodiment, if the residual vibration peak amplitude corresponding to a target fire source location does not exceed the preset vibration threshold, the initial gimbal rotation speed remains unchanged in the polling segment where that location is located; if the residual vibration peak amplitude exceeds the preset vibration threshold, the polling segment where the target fire source location is located is marked as a polling segment to be downgraded. After comparing each target fire source in the raw material storage area, production and processing area, and finished product loading and unloading area, an over-limit judgment result is obtained, marking several polling segments to be downgraded.

[0041] It should be noted that the peak amplitude of residual jitter during the downshifting polling phase is relatively high because the initial gimbal rotation speed is too fast, resulting in excessive inertial torque accumulated during the pause. Reducing the rotation speed can decrease the inertial torque during the pause, thereby reducing the peak amplitude of residual jitter, which forms the basis for downshifting.

[0042] Understandably, a single downshift may not bring the residual vibration peak amplitude back below the preset vibration threshold. Therefore, the downshift processing adopts a closed-loop approach of progressive reduction and re-prediction.

[0043] Specifically, for each polling segment to be downgraded in the over-limit judgment result, the current initial gimbal rotation speed of that polling segment is obtained, and it is reduced by one level according to a preset attenuation ratio to obtain the reduced rotation speed value. The preset attenuation ratio is the decrease in rotation speed of each level relative to the original rotation speed.

[0044] For example, 0.85 is used, meaning that for each reduction level, the rotational speed decreases to 0.85 times its original value. Then, the reduced rotational speed value, along with the weight distribution of each segment of the robotic arm corresponding to that polling segment and the rotational inertia parameters of the two imaging components, is reconstructed into a time sequence and fed into the Long Short-Term Memory (LSTM) network. The network then re-outputs the residual jitter peak amplitude at that rotational speed for the polling segment to be downgraded. Further, the re-output residual jitter peak amplitude is compared again with a preset vibration threshold: if it still exceeds the preset vibration threshold, the rotational speed is reduced by another level according to a preset attenuation ratio, and the signal is fed back into the LSM network to re-output the residual jitter peak amplitude. This process of reduction and re-prediction continues until the residual jitter peak amplitude at a certain level falls below the preset vibration threshold. At this point, the reduction stops, and the rotational speed value corresponding to that level is determined as the target gimbal rotational speed for the polling segment to be downgraded.

[0045] For example, the target fire source in the finished product loading and unloading area of ​​the chemical production area has a high initial gimbal rotation speed due to its far location. Its residual vibration peak amplitude exceeds the preset vibration threshold. After two levels of reduction by 0.85 times, the peak amplitude drops back to below the preset vibration threshold. The rotation speed value after the two levels of reduction is taken as the target gimbal rotation speed of this polling segment.

[0046] Preferably, the above-mentioned step-by-step reduction process is performed on each polling segment to be downgraded to obtain the target gimbal rotation speed corresponding to each polling segment to be downgraded. The polling segments that are not marked use the initial gimbal rotation speed, thereby obtaining the target gimbal rotation speed used by the gimbal to poll each target fire source.

[0047] S104. After the target gimbal rotation speed drives the detection device to reach the target fire source location, it acquires image frame groups, analyzes the edge contours of adjacent frames, determines the degree of image shaking by identifying the degree of matching of the edge contours of adjacent frames, and performs deblurring on frames whose image shaking degree exceeds the preset shaking threshold to obtain a stable image group.

[0048] After the target gimbal rotation drives the detection device to the target fire source location, the infrared and visible light dual-band imaging component is activated, and the fire source image is continuously recorded at fixed acquisition times, resulting in a time-sequential image frame group. The edge contours at the boundary between the flame and the background are extracted from each frame in the image frame group. The edge contours of two adjacent frames in the image frame group are obtained, and the pixel offset of the corresponding point on the same flame edge in the two adjacent frames is calculated to obtain the contour displacement of the two adjacent frames. The degree of matching of the edge contours of the adjacent frames is determined based on the magnitude of the contour displacement; the larger the contour displacement, the lower the degree of matching. The degree of matching determines the image shake of the adjacent frames. For frames whose image shake exceeds a preset shake threshold, Wiener filtering is used to construct a point spread function based on the contour displacement direction and magnitude of the frame to deblur the frame pixel by pixel, resulting in a clear frame after deblurring. The clear frame and the frames whose image shake does not exceed the preset shake threshold are merged according to the acquisition time to obtain a stable image group.

[0049] Specifically, after obtaining the target gimbal rotation speed, the detection device is driven to turn and reach the target fire source location using the target gimbal rotation speed. At this time, the residual vibration peak amplitude after the gimbal stops has been reduced to below the preset vibration threshold, and the shaking degree of the lens image is reduced accordingly, laying the foundation for acquiring clear fire images. After the detection device reaches the target fire source location, the infrared and visible light dual-band imaging components are activated. The infrared imaging component acquires the temperature field distribution of the flame, presenting the thermal radiation intensity of the flame; the visible light imaging component acquires the combustion outline of the flame, presenting the color and shape of the flame. The two imaging components synchronously and continuously acquire fire source images at fixed acquisition times, outputting one frame at each acquisition time, and all frames are arranged in chronological order to form an image frame group.

[0050] It should be noted that even after the rotation speed is reduced, slight residual vibrations remain after the gimbal stops, causing slight image displacement between some frames. Determining which frames are affected by the shaking relies on comparing the edge contours of the flame in adjacent frames. The edge contour refers to the pixel line where the brightness abruptly changes at the boundary between the flame and the background; in visible light frames, this manifests as the light-dark boundary at the outer edge of the flame, and in infrared frames, as the boundary of radiation intensity between the high-temperature area and the low-temperature background. The brightness gradient of adjacent pixels in each frame of the image frame group is detected pixel by pixel. Pixels with brightness gradients exceeding a preset gradient threshold are connected to extract the edge contour of that frame.

[0051] It is understandable that if the image does not shake within the capture interval between two adjacent frames, the outline of the flame edge in the two frames should almost overlap; if shaking occurs, the pixel position of the same flame edge will shift between the two frames.

[0052] Specifically, corresponding points on the same flame edge in two adjacent frames are taken, and the pixel offset of these points between the two frames is calculated to obtain the contour displacement of the two adjacent frames. The larger the contour displacement, the worse the overlap of the edge contours between the two frames, and the lower the degree of matching; the smaller the contour displacement, the higher the degree of matching. The degree of matching is then used to calculate the image shake of the adjacent frames, with a low degree of matching corresponding to a high degree of image shake.

[0053] For example, in the image of a fire source in a raw material storage area, if the pixel offset of corresponding points on the outer edge of the flame in two adjacent frames exceeds 2 pixels, the image shake level is determined to be too high; if the offset is within 1 pixel, the image shake level is determined to be too low. The preset shake threshold is the image shake level corresponding to this allowable offset, which serves as the boundary for determining whether the frame needs deblurring. Further, deblurring is performed on frames whose image shake level exceeds the preset shake threshold. In these frames, the lens shifts at the moment of acquisition, and the flame edge is dragged into blurred stripes along the direction of the shift. The shape of the blur is determined by both the direction and length of the shift. The deblurring process uses Wiener filtering, which is a filtering method that restores a clear image based on the reverse operation of the degradation process. Its restoration effect depends on the characterization of the degradation pattern. The degradation pattern is represented by a point spread function, which describes the shape in which an ideal point light source is scattered after imaging.

[0054] In one embodiment, the extension direction of the blurred stripes is determined based on the contour displacement direction of the frame, and the length of the blurred stripes is determined based on the contour displacement magnitude of the frame. Based on this, a point spread function that matches the actual blurred shape of the frame is constructed. Then, Wiener filtering is used to perform reverse operation on the frame pixel by pixel according to the point spread function to restore the dragged flame edge back to the converged state, thus obtaining the clear frame after deblurring.

[0055] It should be noted that frames whose image shaking does not exceed a preset shaking threshold are already sufficiently clear and are retained directly without deblurring. These clear frames are then merged with the frames that do not exceed the preset shaking threshold according to their acquisition time, restoring a time-ordered frame sequence to obtain a stable image group.

[0056] Preferably, the image frame groups of the infrared band and the visible light band are respectively subjected to the above comparison and deblurring processing to obtain infrared stable image groups and visible light stable image groups, respectively. Both can present a clear flame temperature field and combustion contour, providing clear image support for subsequent identification of the combustion stage of the fire source. In another embodiment, the edge contours of two adjacent frames in the image frame group are obtained. The contour of the previous frame is resampled at equal intervals according to the arc length to obtain a set of reference points. Each reference point is accompanied by its tangential angle and local curvature as a descriptor. On the contour of the next frame, the point with the smallest Euclidean distance and a tangential angle difference of no more than 15 degrees and a relative curvature difference of no more than 0.2 is found as the corresponding point of the same flame edge by the nearest neighbor search method. If the candidate point does not meet the constraints, the reference point is eliminated and does not participate in the calculation. Thus, the pixel offset of the corresponding point of the same flame edge in two adjacent frames is calculated. The median of the offset vectors of all valid corresponding points is taken as the contour displacement d of the two adjacent frames. The direction of d is denoted as θ, and the magnitude is denoted as L. The degree of matching of the edge contours of adjacent frames is determined according to the size of L. The larger L is, the lower the degree of matching. The degree of image shakiness, S, is defined as S = L / L0, where L0 is 1% of the image diagonal length as a normalization benchmark. A frame is considered stable when S does not exceed 0.5, and a shaky frame is identified when S is greater than 0.5, triggering deblurring. For shaky frames, a line segment-like point spread function h is constructed with θ as the blur direction and L as the blur length. Within a square kernel of size L rounded up and plus 1, pixels along a straight line segment passing through the center point and making an angle θ with the horizontal direction are assigned equal weights and normalized to a sum of 1 within the kernel. Pixels outside the line segment are set to zero. h is then substituted into a Wiener filter to deblur the frame pixel by pixel. An empirically chosen noise-to-signal ratio of 0.01 is used to obtain a clear frame after deblurring. The clear frame and the stable frame are then merged according to the acquisition time to obtain a stable image group.

[0057] S105. Segment the flame region from the stable image group, extract the flame combustion area and temperature distribution, identify the current combustion stage of the fire source through support vector machine, calculate the minimum number of data frames required for reliable updates of the combustion stage, and determine the observation window duration of the current target fire source location.

[0058] From the stable image set, the closed contour of the flame-background boundary is extracted from the visible light frame to delineate the flame region. The total number of pixels within the flame region is counted to calculate the flame combustion area. For the infrared frame, the radiation intensity of each pixel within the flame region is read to calculate the temperature distribution. The statistics of the flame combustion area and temperature distribution are concatenated to form a feature vector characterizing the current combustion state of the fire source. The feature vector is input into a pre-established support vector machine. The support vector machine uses the feature vectors of historical fire sources in the initial combustion, development, and intense combustion stages as training samples to obtain support vectors and classification hyperplanes for each combustion stage. The current combustion stage of the fire source is determined based on the region corresponding to each combustion stage that the feature vector falls into. The fluctuation amplitude of the combustion stage between adjacent discrimination results is obtained. The number of consecutive frames corresponding to the fluctuation amplitude converging to within a preset fluctuation threshold is determined as the minimum number of data frames required for reliable updates of the combustion stage. The observation window duration for the current target fire source location is determined by dividing the minimum number of data frames by the number of frames acquired per second by the multi-band imaging component.

[0059] After obtaining a stable set of images, the combustion stage of the current target fire source is determined, and based on this, it is determined how long the location of the target fire source must remain to acquire a sufficient number of clear frames. The flame morphology and thermal radiation intensity of the fire source differ significantly in the initial combustion, development, and intense combustion stages. Determining the combustion stage relies on the extraction of objective features of the flame area.

[0060] Specifically, the flame region refers to the set of pixels in the image that truly belong to the burning flame. From the visible light frames of the stable image group, closed contours with abrupt brightness changes at the boundary between the flame and the background are extracted, and the set of pixels enclosed by these closed contours is defined as the flame region. This definition eliminates interference from the background and smoke plumes, ensuring that subsequent feature extraction focuses solely on the flame itself. In one embodiment, two types of objective features are extracted from the flame region. The first is the flame burning area. The total number of pixels within the flame region is counted, and the total number of pixels is converted into the actual burning area based on the field of view of the imaging component and the distance to the target fire source. The burning area reflects the spatial spread of the fire. The second is the temperature distribution. For the infrared frames of the stable image group, the radiation intensity of each pixel within the flame region is read, and the temperature distribution of the flame region is calculated based on the correspondence between radiation intensity and surface temperature. The temperature distribution reflects the heat gradient between the high-temperature core and the outer edge of the flame.

[0061] It should be noted that a single feature is insufficient to distinguish the combustion stages. In the initial combustion stage, the combustion area is small and the temperature distribution is concentrated in a localized area; in the development stage, the combustion area expands and the high-temperature region extends outward; in the intense stage, the combustion area is large and the overall temperature distribution rises. Therefore, the statistical quantities of the flame combustion area and temperature distribution, including the mean of the temperature distribution and the proportion of high-temperature pixels, are concatenated in a fixed order to form a feature vector representing the current combustion state of the fire source. Each dimension of the feature vector corresponds to an objective feature.

[0062] Understandably, the correspondence between feature vectors and combustion stages is determined using a support vector machine (SVM). SVM is a machine learning algorithm used for classification. Its basic principle is to find a classification hyperplane in the feature space that separates samples of different classes as much as possible. The sample points closest to the hyperplane are called support vectors, and the support vectors determine the position and orientation of the hyperplane.

[0063] In one embodiment, the training of the support vector machine (SVM) relies on historical fire source data. Visible and infrared frames of historical fire sources at each of the initial combustion, development, and intense combustion stages are collected. The flame area and temperature distribution of each frame are extracted as described above to form feature vectors, which are then labeled with their corresponding combustion stages. These feature vectors are input into the SVM as training samples to obtain the support vectors and classification hyperplanes that divide each combustion stage. For the initial combustion, development, and intense combustion stages, multiple classification hyperplanes are constructed in a pairwise manner, collectively dividing the feature space into regions corresponding to the three combustion stages. During discrimination, the feature vector of the current fire source is fed into the pre-built SVM, and the combustion stage of the current fire source is determined based on which region the feature vector falls into.

[0064] For example, in a chemical production area, if the ignition source in the raw material storage area has a small combustion area and concentrated temperature distribution in a certain frame of its feature vector, falling into the region corresponding to the initial combustion stage, then the judgment result for that frame is the initial combustion stage. Furthermore, the judgment result of a single frame is easily affected by residual slight shaking and the flickering of the flame itself, so the combustion stage can only be reliably updated after the judgment results of multiple consecutive frames are stable and consistent.

[0065] Specifically, the stable image group is fed into a support vector machine frame by frame to obtain the combustion stage discrimination result frame by frame, and the fluctuation amplitude between adjacent discrimination results is obtained, that is, whether the discrimination results of adjacent frames are consistent and the severity of the jump between the three combustion stages; when the discrimination results of several consecutive frames no longer jump and the fluctuation amplitude converges to within the preset fluctuation threshold, the number of consecutive frames is determined as the minimum number of data frames required for reliable update of the combustion stage.

[0066] For example, if the preset fluctuation threshold is set to five consecutive frames with completely consistent discrimination results, then the minimum number of data frames is set to these five frames. Finally, the observation window duration is determined. The multi-band imaging component acquires a fixed number of frames per second; this number is the sampling frequency. The observation window duration must ensure that the number of frames acquired within this duration is not less than the minimum number of data frames; therefore, it is calculated by dividing the minimum number of data frames by the sampling frequency. Symbolically, let the minimum number of data frames be n, the sampling frequency be f frames per second, and the observation window duration be t seconds, then t = n ÷ f, where n is the minimum number of data frames required for reliable updates during the combustion phase, and f is the number of frames acquired per second by the multi-band imaging component. For example, when the minimum number of data frames is 5 frames and the sampling frequency is 10 frames per second, the observation window duration is 0.5 seconds.

[0067] Preferably, feature vectors are extracted from each target fire source in the raw material storage area, production and processing area, and finished product loading and unloading area, and the combustion stage is determined. The observation window duration for each target fire source location is then determined. For fire sources with more intense combustion and more intense flame flickering, the minimum number of data frames required for the fluctuation and convergence of the discrimination result is larger, and the corresponding observation window duration is longer. This causes the time the gimbal stays at each target fire source location to vary depending on the difficulty of discerning the current combustion stage. In another implementation, for visible light frames from the stable image set, median filtering is first used to suppress speckle noise. Then, the Otsu method is used to adaptively calculate the segmentation threshold from the grayscale histogram to generate a binary mask for the flame foreground. Canny edge detection is performed on the mask to obtain an edge point set. The flame region is delineated by connecting the breakpoints through morphological closing operations of 3*3 structuring elements and forming a closed contour. The total number of pixels within the flame region is counted to calculate the flame burning area. For infrared frames, the radiation intensity of each pixel within the flame region is read to calculate the temperature distribution. The flame burning area is concatenated with the mean, variance, and maximum value of the temperature distribution to form a feature vector representing the current combustion state of the fire source. The feature vector is input into a pre-established support vector machine. The support vector machine uses the feature vectors of historical fire sources in the initial combustion, development, and intense combustion stages as training samples to obtain support vectors and classification hyperplanes for each combustion stage. The current combustion stage of the fire source is determined based on the region corresponding to each combustion stage that the feature vector falls into. A sliding window of length N is used to extract the combustion stage discrimination results of the most recent N frames. The number of times k is inconsistent between adjacent frames within the window is counted, and the fluctuation amplitude P is recorded as k / (N-1). The larger P is, the more frequently the stage discrimination jumps within the window. The preset fluctuation threshold is set to 0.1. When P is continuously kept below 0.1, the discrimination is considered to have converged. At this time, the number of consecutive frames corresponding to the window is the minimum number of data frames required for reliable update of the combustion stage. For example, N is 10 and k is reduced to 0 to meet the condition, and the minimum number of data frames is 10. The minimum number of data frames is divided by the number of frames collected per second by the multi-band imaging component to determine the observation window duration of the current target fire source position.

[0068] S106. Within the observation window duration, continuously acquire stable image frames. After accumulating to the minimum number of data frames, extract the centroid pixel coordinates of the flame outline in each frame. Map the centroid pixel coordinates to the three-dimensional coordinate system of the monitoring area through coordinate back projection to obtain the three-dimensional coordinates of the current target fire source.

[0069] Within the observation window, stable image frames are continuously acquired. A count is accumulated for each acquired stable image frame. If the accumulated count does not reach the minimum number of data frames, acquisition continues; if the accumulated count reaches the minimum number of data frames, acquisition stops, resulting in a stable frame group consisting of the minimum number of data frames. For each frame in the stable frame group, the flame outline at the boundary between the flame and the background is extracted. The geometric center of the pixels enclosed by the flame outline is calculated to obtain the centroid pixel coordinates of the flame outline in that frame. The centroid pixel coordinates of each frame are arranged in frame order, and their average value is taken to obtain the stable centroid pixel coordinates of the current target fire source. The pre-calibrated camera parameters of the detection device and the depth distance of the flame area read in the infrared band are obtained. Coordinate back projection is used to map the stable centroid pixel coordinates from the two-dimensional image plane to the three-dimensional coordinate system of the monitoring area based on the camera parameters and the depth distance, obtaining the precise three-dimensional coordinates of the current target fire source.

[0070] After determining the observation window duration for the current target fire source location, stable image frames are continuously acquired for the target fire source within this observation window duration. The stable image frame is formed by merging the previously deblurred clear frame with frames that do not exceed a preset shaking threshold. The degree of image shaking has been suppressed to a usable level, laying the foundation for extracting the flame location.

[0071] Specifically, the acquisition process includes a counter. The counter increments by 1 for each stable image frame acquired. If the accumulated count does not reach the minimum number of data frames, acquisition continues; if the accumulated count reaches the minimum number of data frames, acquisition stops, resulting in a stable frame set consisting of the minimum number of data frames. The minimum number of data frames is the number of frames required for reliable updates during the aforementioned combustion phase; therefore, the accumulated stable frame set at the time of acquisition cessation precisely meets the sample size for reliable positioning.

[0072] It should be noted that the location of the fire source is represented by the centroid of the flame outline. As the flame continues to oscillate during combustion, the center of the flame outline in a single frame fluctuates, making it difficult to determine a stable fire source location based on a single frame. Therefore, the centroids of multiple frames in a stable frame group are extracted and averaged. In one embodiment, for each frame in the stable frame group, the flame outline at the boundary between the flame and the background is extracted. This flame outline is a closed pixel line connecting the flame's outer edge, where brightness changes abruptly. The geometric center of the pixels enclosed by the flame outline is calculated. The x-coordinate of the geometric center is the average of the x-coordinates of all pixels within the outline, and the y-coordinate is the average of the y-coordinates of all pixels, yielding the centroid pixel coordinates of the flame outline for that frame.

[0073] Understandably, the centroid pixel coordinates obtained frame by frame fluctuate within a small range due to flame swaying. The centroid pixel coordinates of each frame in the stable frame group are arranged in frame order, and their average value is taken. The average horizontal and vertical coordinates together constitute the stable centroid pixel coordinates of the current target fire source. These stable centroid pixel coordinates smooth out single-frame jitter and characterize the stable position of the fire source on the image plane within the observation window. Furthermore, these stable centroid pixel coordinates are mapped from the two-dimensional image plane to the three-dimensional coordinate system of the monitoring area. The stable centroid pixel coordinates only indicate the position of the fire source on the screen and do not contain information about the distance between the fire source and the detection device; therefore, the mapping relies on two pre-acquired data points: camera parameters and depth distance.

[0074] Specifically, camera parameters refer to the internal and external parameters of the imaging component of the detection equipment, which are pre-calibrated. Internal parameters include focal length and principal point position on the image plane, describing the geometric correspondence between pixel coordinates and the optical axis of the imaging component. External parameters include the installation position and orientation of the imaging component in the three-dimensional coordinate system of the monitoring area, obtained through a one-time calibration during equipment installation. Depth distance refers to the distance from the fire source along the optical axis of the imaging component to the imaging component, actively measured by a pulsed laser ranging unit coaxially mounted with the imaging component. The ranging unit illuminates the target area with a 905 nm laser pulse and calculates the depth distance d by measuring the flight time t between the emitted and echo pulses. The calculation relationship is d = c * t / 2, where c is the speed of light. This method does not rely on the radiation characteristics of the flame itself and is unaffected by factors such as flame temperature differences, emissivity fluctuations, atmospheric attenuation, and smoke plume obstruction. It can stably output depth measurement results with an accuracy better than 0.5 meters, meeting the accuracy requirements for depth input in subsequent coordinate back projection. The laser ranging unit and the imaging component share the same installation reference, and the coordinate system offset between the two is compensated during the factory calibration stage, without the need for online calculation of coordinate system transformation.

[0075] In one embodiment, coordinate back-projection is used to complete the mapping. Coordinate back-projection is the inverse operation of pinhole imaging: pinhole imaging projects a three-dimensional point into two-dimensional pixels, while back-projection, based on the focal length and principal point position in the camera parameters, restores the stable centroid pixel coordinates into a ray originating from the optical center of the imaging component and passing through that pixel; then, the specific position of the fire source on this ray is determined by the depth distance, obtaining the three-dimensional coordinates of the fire source relative to the imaging component; finally, the three-dimensional coordinates are transformed to the three-dimensional coordinate system of the monitoring area using the installation position and orientation in the camera's external parameters, obtaining the precise three-dimensional coordinates of the current target fire source. These three-dimensional coordinates indicate the specific spatial location of the fire source in the raw material storage area, production and processing area, or finished product loading and unloading area with three components: lateral position, longitudinal position, and height.

[0076] For example, in the finished product loading and unloading area of ​​a chemical production zone, the stable centroid pixel coordinates of a target fire source are located in the lower right of the screen. The depth distance read by the infrared band indicates that the fire source is far from the detection device. After the coordinate back projection restores the ray and locates it along the ray according to the depth distance, it is transformed to the three-dimensional coordinate system of the monitoring area to obtain the precise three-dimensional coordinates of the fire source above the cargo position on the east side of the loading and unloading area.

[0077] Preferably, stable frame groups are collected for each target fire source identified in the raw material storage area, production and processing area, and finished product loading and unloading area. Stable centroid pixel coordinates are extracted, and coordinate back projection is performed to obtain the precise three-dimensional coordinates of each target fire source. The precise three-dimensional coordinates of each target fire source collectively indicate the spatial distribution of multiple fire points within the monitoring area, enabling fire rescue teams to locate the actual position of each fire source based on these precise three-dimensional coordinates.

[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for accurate detection of multispectral fire source location in complex fire scenarios, characterized in that, The method includes: The infrared imaging component mounted on the PTZ camera scans the entire monitoring area, identifies areas with abnormal temperatures, and locates the range of each target fire source, thereby setting the initial PTZ camera rotation speed. The initial gimbal rotation speed and the mass distribution data of the detection device are input into a pre-established long short-term memory network to obtain the predicted mechanical vibration amplitude when the detection device reaches the target fire source location. When the predicted mechanical vibration amplitude exceeds the preset vibration threshold, the initial gimbal rotation speed is downgraded according to the preset attenuation ratio to obtain the target gimbal rotation speed. After the target gimbal rotation speed drives the detection device to reach the target fire source location, it acquires image frame groups, analyzes the edge contours of adjacent frames, determines the degree of image shaking by identifying the degree of matching of the edge contours of adjacent frames, and performs deblurring on frames whose image shaking degree exceeds the preset shaking threshold to obtain a stable image group. The flame region is segmented from the stable image group, the flame combustion area and temperature distribution are extracted, the current combustion stage of the fire source is identified by support vector machine, the minimum number of data frames required to update the combustion stage is calculated, and the observation window duration of the current target fire source location is determined. Stable image frames are continuously acquired within the observation window duration. After accumulating to the minimum number of data frames, the centroid pixel coordinates of the flame outline in each frame are extracted. The centroid pixel coordinates are then mapped to the three-dimensional coordinate system of the monitoring area through coordinate back projection to obtain the three-dimensional coordinates of the current target fire source.

2. The method for accurate detection of multispectral fire source location in complex fire scenarios according to claim 1, characterized in that, The process involves using an infrared imaging component mounted on a pan-tilt unit to perform a full-range scan of the monitored area, identify areas of abnormal temperature, and locate the range of each target fire source. Based on this, the initial pan-tilt rotation speed is set, including: The monitoring area is scanned line by line by the infrared imaging component mounted on the pan-tilt unit to obtain the radiation intensity of each pixel. Continuous pixels with radiation intensity higher than a preset abnormal threshold are clustered to obtain the temperature abnormal area. The temperature anomaly area is extended outward along the direction of radiation intensity from high to low until the radiation intensity drops back to the background radiation intensity, thus defining the outline boundary of the range of each target fire source. Based on the outline boundary of each target fire source range, extract the centroid azimuth and elevation angle of each fire source, and determine the polling order of the gimbal to arrive at each target fire source in sequence according to the azimuth angle. Based on the distance between two adjacent target fire sources in the polling order, a higher speed gear is matched for polling segments with a distance greater than a preset distance threshold, and a lower speed gear is matched for polling segments with a distance less than the preset distance threshold, so as to obtain the initial gimbal speed of each polling segment.

3. The method for accurate detection of multispectral fire source location in complex fire scenarios according to claim 2, characterized in that, The step of extracting the centroid azimuth and elevation angles of each fire source and determining the polling order of the gimbal to arrive at each target fire source in sequence according to the azimuth angles includes: calculating the azimuth distance span between adjacent fire sources for each pair of target fire sources based on their spatial distribution, arranging each target fire source in ascending order of azimuth angles, and determining the polling order of the gimbal to arrive at each target fire source in sequence.

4. The method for accurate detection of multispectral fire source location in complex fire scenarios according to claim 1, characterized in that, The step of inputting the initial gimbal rotation speed and the mass distribution data of the detection device into a pre-established long short-term memory network to obtain the predicted mechanical vibration amplitude when the detection device reaches the target fire source location includes: Based on the mass distribution data, the weight distribution of each segment of the gimbal robotic arm along the arm length direction and the rotational inertia parameters of the imaging component at the installation position are extracted. The initial gimbal rotation speed, weight distribution and rotational inertia parameters are arranged according to the continuous rotation time before the gimbal reaches the target fire source position to obtain a time sequence reflecting the rotation process. The residual flutter waveforms of the detection equipment after reaching the target point with different initial gimbal rotation speeds and different mass distributions during its historical operation are collected. The vibration amplitude decay curve from the moment of pause is extracted. The time series is used as input and the peak amplitude of the vibration amplitude decay curve is used as output to label training samples and iteratively update the weights of the internal memory units of the long short-term memory network. The time sequence is fed into the pre-established long short-term memory network step by step, and the peak amplitude of residual flutter is recursively output based on the rotational inertia parameter and speed change at the pause time to determine the mechanical vibration amplitude prediction result.

5. The method for accurate detection of multispectral fire source location in complex fire scenarios according to claim 1, characterized in that, When the predicted mechanical vibration amplitude exceeds a preset vibration threshold, the initial gimbal rotation speed is downgraded according to a preset attenuation ratio to obtain the target gimbal rotation speed. This includes: comparing the residual flutter peak amplitudes in the predicted mechanical vibration amplitude with the preset vibration threshold one by one; marking the polling segment where the target fire source location exceeds the preset vibration threshold as the polling segment to be downgraded; reducing the initial gimbal rotation speed of the polling segment step by step according to the preset attenuation ratio; and re-outputting the residual flutter peak amplitude by the long short-term memory network for each reduction until it falls below the preset vibration threshold to obtain the target gimbal rotation speed.

6. The method for accurate detection of multispectral fire source location in complex fire scenarios according to claim 1, characterized in that, After the target gimbal rotation speed drives the detection device to reach the target fire source location, it acquires a group of image frames, analyzes the edge contours of adjacent frames, determines the degree of image shake by identifying the degree of matching of the edge contours of adjacent frames, and deblurs frames whose image shake exceeds a preset shake threshold to obtain a stable image group, including: Activate the infrared and visible light dual-band imaging component, continuously record fire source images at fixed acquisition times, and extract the edge contours at the boundary between the flame and the background for each frame. The pixel offset is calculated for the corresponding point of the same flame edge in two adjacent frames to obtain the contour displacement of the two adjacent frames. The degree of matching of the edge contour of the adjacent frames is determined based on the magnitude of the contour displacement, and the degree of image shaking is determined based on the degree of matching. For frames whose image shake exceeds a preset shake threshold, Wiener filtering is used to construct a point spread function based on the frame contour displacement direction and displacement magnitude to deblur each pixel. The clear frames after deblurring are then merged with the frames that do not exceed the preset shake threshold according to the acquisition time to obtain a stable image group.

7. The method for accurate detection of multispectral fire source location in complex fire scenarios according to claim 1, characterized in that, The process of segmenting the flame region from the stable image group, extracting the flame combustion area and temperature distribution, identifying the current fire source combustion stage using a support vector machine, calculating the minimum number of data frames required for updating the combustion stage, and determining the observation window duration for the current target fire source location includes: extracting closed contours to delineate the flame region for visible light band frames, calculating the flame combustion area by counting the total number of pixels, and calculating the temperature distribution by reading the radiation intensity of each pixel within the flame region for infrared band frames; concatenating the statistics of the flame combustion area and temperature distribution to form a feature vector; inputting the feature vector into the support vector machine, determining the combustion stage based on the region corresponding to each combustion stage; determining the minimum number of data frames by converging the fluctuation amplitude of adjacent discrimination results of combustion stages to within a preset fluctuation threshold, and dividing this number by the number of frames acquired per second by the imaging component to determine the observation window duration.

8. The method for accurate detection of multispectral fire source location in complex fire scenarios according to claim 7, characterized in that, The step of inputting the feature vector into the support vector machine and determining the combustion stage based on the region corresponding to each combustion stage includes: the support vector machine uses the feature vectors of historical fire sources in the initial combustion, development and intense combustion stages as training samples to obtain the support vectors and classification hyperplanes that divide each combustion stage, and determines the combustion stage of the current fire source based on the region where the feature vector falls.

9. The method for accurate detection of multispectral fire source location in complex fire scenarios according to claim 1, characterized in that, The process of continuously acquiring stable image frames within the observation window duration, accumulating to reach the minimum number of data frames, extracting the centroid pixel coordinates of the flame outline in each frame, and mapping the centroid pixel coordinates to the three-dimensional coordinate system of the monitoring area through coordinate back projection to obtain the three-dimensional coordinates of the current target fire source includes: accumulating and counting stable image frames until the minimum number of data frames is reached to obtain a stable frame group; extracting the flame outline for each frame, calculating the geometric center of the surrounding pixels to obtain the centroid pixel coordinates, and taking the average value according to the frame order to obtain the stable centroid pixel coordinates; obtaining the camera parameters and the flame area depth distance read by the infrared band, and mapping the stable centroid pixel coordinates to the three-dimensional coordinate system of the monitoring area using coordinate back projection to obtain the three-dimensional coordinates.