Intelligent fire monitoring remote monitoring method based on image processing
Patent Information
- Application Number
- CN202610556292.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-24
- Publication Date
- 2026-08-18
AI Technical Summary
[0002]通过热成像设备对场所进行远程火灾监测,相机的稳定成像是算法可靠工作的基础;然而,热成像设备在实际运行中,为保障测温准确性与图像质量,会内置一系列自校准与自适应处理机制;其中,周期性的快门校正用于补偿探测器像元的响应漂移与非均匀性;同时,相机内部的自动增益控制与动态范围压缩算法也会根据场景温度分布实时调整输出信号的映射关系;这些必要的内部处理,往往导致相机输出视频在瞬间发生全局性变化,例如整体亮度跳变、对比度重置、短暂帧冻结或灰度映射曲线畸变,致使校正前后的连续帧之间失去直接的像素值可比性
1、本发明通过识别校正干扰区间并构建受约束的单调映射函数进行灰度调整,恢复了视频序列的时序可比性,这使得后续的时序分析算法能够基于连续、一致的图像数据进行,从根本上避免了因数据源突变而引发的分析中断;通过引入设备成像噪声指纹模型筛选出可靠的背景统计区域,并采用对灰度映射变化鲁棒的火情特征进行跨校正干扰区间的连续证据累计,既从数据源头降低了噪声干扰,又通过鲁棒特征的设计减少了对绝对灰度值的依赖,更通过巧妙的证据累计机制保护了火情证据链在校正期间不被切断或污染,从而有效防止了漏报和误抑制;通过受约束的函数估计和基于估计可信度的降级检测策略双重保障,实现了安全、自适应的处理,当补偿可信度高时,能有效恢复图像一致性;当可信度不足时,自动切换至基于鲁棒特征和前后证据融合的降级模式,确保了在任何情况下系统都能以最安全可靠的方式运行,最大限度地降低了误报风险;通过输出包含校正干扰区间标记、估计可信度等在内的图像质量标记,使远程监控人员能够清晰了解报警发生时的系统状态和数据质量,为高效、准确的复核与决策提供了关键支持。
Smart Images

Figure CN122598352A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of fire monitoring, and in particular to an intelligent remote monitoring method for fire monitoring based on image processing. Background Technology
[0002] For remote fire monitoring using thermal imaging equipment, stable camera imaging is fundamental to the reliable operation of the algorithm. However, in actual operation, thermal imaging equipment incorporates a series of self-calibration and adaptive processing mechanisms to ensure temperature measurement accuracy and image quality. Among these mechanisms, periodic shutter correction is used to compensate for the response drift and non-uniformity of detector pixels. Simultaneously, the camera's internal automatic gain control and dynamic range compression algorithms adjust the mapping relationship of the output signal in real time according to the scene's temperature distribution. These necessary internal processes often cause instantaneous global changes in the camera's output video, such as overall brightness jumps, contrast resets, brief frame freezes, or grayscale mapping curve distortions, resulting in a loss of direct pixel value comparability between consecutive frames before and after correction.
[0003] In existing technologies, mainstream fire detection algorithms heavily rely on the temporal continuity and grayscale stability of video sequences. Whether it is motion and change detection based on inter-frame difference, scene anomaly perception based on background modeling, or early fire judgment based on temporal consistency accumulation, their effectiveness is based on the assumption that the grayscale relationship between adjacent frames can be continuously compared. When events such as shutter correction occur, the algorithm is very likely to misidentify the global mapping mutations introduced by the camera itself as meaningless changes in illumination or global interference in the scene, thus actively suppressing alarms. In addition, if distorted frame data is used to update the background model, it will pollute the statistical characteristics of the model, resulting in the inability to accurately perceive the real fire situation for a long period of time after the correction is completed, causing missed alarms or alarm delays. Summary of the Invention
[0004] To address the technical problems existing in the background art, this invention proposes an intelligent remote fire monitoring method based on image processing, the specific solution of which is as follows: An intelligent fire monitoring and remote control method based on image processing includes the following steps: S1. Acquire consecutive image frames of thermal imaging video, construct a thermal imaging video sequence, and identify the correction interference interval where the image grayscale mapping changes abruptly based on the thermal imaging video sequence. S2. Pre-build a device imaging noise fingerprint model, and select a background statistical region for each video frame based on the device imaging noise fingerprint model; S3. When it is determined that the current frame is in the correction interference interval, the latest stable video frame before the start of the correction interference interval is determined as the reference frame; based on the gray-scale statistical relationship between the background statistical region of the current frame and the reference frame, a constrained monotonic mapping function is constructed, and the gray-scale mapping of the current frame is adjusted through the monotonic mapping function to generate a video frame with consistent gray-scale across the correction interference interval. S4. Based on the grayscale-uniformed video frames, extract fire features that are robust to grayscale mapping changes, and perform continuous evidence accumulation and fire determination across the correction interference interval. S5. When the estimation confidence of the monotonic mapping function is insufficient, a downgrade detection strategy is enabled. The downgrade detection strategy includes: skipping the grayscale mapping adjustment of the current frame and directly determining the fire situation based on the fire features and fusing the separation evidence before and after the correction interference interval. S6. Output the fire situation determination result and associated image quality markers.
[0005] Furthermore, in S1, the correction interference range for sudden changes in image grayscale mapping based on thermal imaging video sequence identification is as follows: Calculate the overall similarity abruptness index between the current frame image and the previous frame image; Calculate the grayscale distribution difference index between the current frame image and the previous frame image; Detect frame freeze indicators in a video sequence; Detect abrupt changes in fixed-pattern noise energy in an image and generate a fixed-noise energy abrupt change index; Based on the overall similarity mutation index, the grayscale distribution difference index, the frame freeze index, and the fixed noise energy mutation index, a correction interference confidence level is generated. When the confidence level of the correction interference exceeds a preset first threshold, it is determined that the frame has entered the correction interference interval.
[0006] Furthermore, in S2, constructing the device imaging noise fingerprint model includes: During the preset fire-free period, for each pixel, the frequency of abnormally high values, the frequency of abnormal gray-level jumps, and the degree of gray-level consistency deviation with neighboring pixels are counted. Based on the abnormal high frequency, the abnormal gray-level jump frequency, and the gray-level consistency deviation, the probability value of bad pixels for each pixel is calculated, and a bad pixel probability map is generated; wherein, the bad pixel probability map constitutes the imaging noise fingerprint model of the device.
[0007] Furthermore, in S2, a background statistical region is selected for each video frame based on the device imaging noise fingerprint model, as follows: Based on the bad pixel probability map, pixels whose bad pixel probability value exceeds the preset second threshold are excluded; Pixels within the moving region are excluded based on inter-frame difference information; Exclude locally highlighted areas and their morphologically dilated neighboring pixels; From the remaining pixels, the region with the smallest gray-level variance within a short time window is selected as the background statistical region.
[0008] Furthermore, in S3, video frames with uniform grayscale across the correction interference interval are generated, as follows: Extract a set of grayscale quantiles with the same percentile from the background statistical regions of the current frame and the reference frame respectively as anchor points; Using the anchor point of the current frame as input and the anchor point of the reference frame as the target, a constrained piecewise linear monotonic function is constructed as a monotonic mapping function. Calculate the fitting residual of the monotonic mapping function and evaluate the reliability of the uniformization process by combining it with the coverage of the background statistical region. When the confidence level is higher than the preset third threshold, the monotonic mapping function is applied to perform grayscale transformation on the current frame to generate a grayscale-consistent video frame.
[0009] Furthermore, in S4, based on the grayscale-uniformed video frames, robust fire features to grayscale mapping changes are extracted as follows: The extracted fire features that are robust to grayscale mapping changes include: The difference between the high grayscale quantile and the median grayscale quantile within the candidate region; The difference between the average gray level of the candidate region and the average gray level of its surrounding annular background region; The trend of changes in the area, perimeter, or compactness of candidate hotspot connected regions over time.
[0010] Furthermore, in S4, continuous evidence accumulation and fire situation determination are performed across the aforementioned correction interference interval, as follows: Within the interference correction interval, pause the iteration of the background statistical model used for fire determination and disable evidence based on inter-frame differences; The fire evidence score is accumulated in the first time window before the start of the correction interference interval and the second time window after the end of the correction interference interval, respectively. A fusion judgment is made based on the accumulated fire evidence scores within the first and second time windows.
[0011] Furthermore, the scores for fire evidence are accumulated separately as follows: For the video frame sequence within the first time window, the first evidence score is accumulated using a sliding window method; For the video frame sequence within the second time window, the second evidence score is accumulated using a sliding window method.
[0012] Furthermore, the scores for fire evidence are accumulated separately as follows: The first evidence score and the second evidence score are weighted and summed, or the difference in the growth of the second evidence score relative to the first evidence score is compared, to obtain the final basis for determining the fire situation.
[0013] Furthermore, in S5, the degradation detection strategy also includes: When the confidence level of the estimate is insufficient, the location information of the suspected fire area identified before the correction interference interval is retained, and the suspected fire area is re-verified after the correction interference interval ends.
[0014] Compared with the prior art, the present invention can achieve at least the following beneficial effects: 1. This invention restores the temporal comparability of video sequences by identifying and correcting interference intervals and constructing a constrained monotonic mapping function for grayscale adjustment. This allows subsequent temporal analysis algorithms to be based on continuous and consistent image data, fundamentally avoiding analysis interruptions caused by sudden changes in data sources. By introducing a device imaging noise fingerprint model to screen reliable background statistical regions, and employing robust fire features to accumulate continuous evidence across correction interference intervals, this invention reduces noise interference at the data source, reduces dependence on absolute grayscale values through robust feature design, and protects the fire evidence chain from being interrupted or obscured during correction through a clever evidence accumulation mechanism. The system effectively prevents missed alarms and false suppressions by employing a dual-protection approach: constrained function estimation and a degradation detection strategy based on estimation confidence. This ensures safe and adaptive processing, effectively restoring image consistency when compensation confidence is high, and automatically switching to a degradation mode based on robust features and fusion of prior and subsequent evidence when confidence is insufficient. This ensures the system operates in the safest and most reliable manner under any circumstances, minimizing the risk of false alarms. By outputting image quality markers including correction interference interval labels and estimation confidence, remote monitoring personnel can clearly understand the system status and data quality at the time of alarm occurrence, providing crucial support for efficient and accurate review and decision-making.
[0015] 2. This invention identifies correction interference intervals caused by the internal processing of thermal imaging cameras and performs frame consistency processing within these intervals based on a constrained grayscale mapping function estimated from a reliable background region, effectively restoring the temporal comparability of video sequences. Simultaneously, by combining device noise fingerprint modeling, robust fire feature extraction for grayscale mapping changes, and an evidence protection accumulation mechanism across interference intervals, the robustness and continuity of the algorithm during camera self-calibration are significantly improved, avoiding false suppression, missed detections, and model contamination caused by inter-frame abrupt changes in traditional methods. Furthermore, with a degradation strategy based on credibility assessment and transparent output including quality markers, a safe and reliable adaptive processing and maintainable decision-making closed loop are achieved, thus solving the problem of fire monitoring interruption and reliability degradation caused by periodic camera calibration. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein: Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation
[0017] Embodiments of the present invention are described in detail below. Examples of these embodiments are illustrated in the accompanying drawings, wherein the same or similar symbols denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0018] Please refer to Figure 1 This invention provides an intelligent remote fire monitoring method based on image processing, comprising the following steps: S1. Acquire consecutive image frames of thermal imaging video, construct a thermal imaging video sequence, and identify correction interference intervals where abrupt changes occur in the image grayscale mapping based on the thermal imaging video sequence.
[0019] It should be noted that acquiring consecutive image frames of thermal imaging video forms a thermal imaging image sequence arranged in chronological order; the correction interference interval specifically refers to the brief period of time during which the grayscale mapping relationship of the output image undergoes a global, non-scene-induced abrupt change due to internal processing such as periodic shutter correction, automatic gain adjustment, or dynamic range compression performed by the thermal imaging camera. Within this interval, the grayscale values of consecutive frames are no longer directly comparable.
[0020] In an optional embodiment, in S1, the correction interference interval where the image grayscale mapping abruptly changes is identified based on the thermal imaging video sequence, as follows: Calculate the overall similarity abruptness index between the current frame image and the previous frame image; It should be noted that the overall similarity mutation index is used to quantify the changes in the overall structure of two adjacent image frames, and its value decreases sharply when a global mapping mutation occurs. This index can be obtained by calculating the structural similarity index or cross-correlation coefficient of two image frames and observing their difference relative to historical stationary periods.
[0021] Calculate the grayscale distribution difference index between the current frame image and the previous frame image; It should be noted that the gray-level distribution difference index is used to quantify the difference in global gray-level statistical characteristics between two adjacent frames of images. Specifically, it can be obtained by calculating the Bach distance between the gray-level histograms of the two frames, the complement of the histogram intersection-union ratio, or by comparing the difference in their cumulative distribution functions at specific quantile points (such as the 10th, 50th, and 90th percentiles).
[0022] Detect frame freeze indicators in a video sequence; It should be noted that the frame freeze index is used to detect the phenomenon of temporary stagnation in the output image caused by camera processing. Specifically, it can be calculated by summing the absolute differences between multiple consecutive frames (e.g., 3-5 frames). If this sum is consistently below a very low noise threshold, but anomalies are found when combined with the statistical characteristics of image noise (e.g., local variance), it may be determined as a frozen frame rather than a true static scene.
[0023] Detect abrupt changes in fixed-pattern noise energy in an image and generate a fixed-noise energy abrupt change index; It should be noted that abrupt changes in the energy of fixed-pattern noise refer to whether the energy of fixed-pattern noise such as fixed stripes, grids, or spots caused by detector non-uniformity in the detected image changes abruptly. This typically occurs before or after the calibration operation. Abrupt changes can be detected by learning a fixed noise template during the calibration phase and then calculating the residual energy of the current frame versus the template in the frequency domain (e.g., observing the energy of specific frequency components via Fourier transform) or the spatial domain during runtime.
[0024] Based on the overall similarity mutation index, the grayscale distribution difference index, the frame freeze index, and the fixed noise energy mutation index, a correction interference confidence level is generated. When the confidence level of the correction interference exceeds a preset first threshold, it is determined that the frame has entered the correction interference interval.
[0025] It should be noted that generating the correction interference confidence score involves normalizing the overall similarity mutation index, grayscale distribution difference index, frame freeze index, and fixed noise energy mutation index obtained above, and then fusing them through a weighted summation or logical judgment rule. For example, a sub-threshold can be set for each index. When the number of indices exceeding the sub-threshold reaches a certain number, or the weighted total score exceeds a comprehensive threshold (i.e., a preset first threshold), it is determined that the frame has entered the correction interference interval. This preset first threshold can be empirically determined by statistically analyzing the index distribution of a large number of normal frames and known correction event frames in typical scenarios.
[0026] It should be noted that the preset first threshold is used to determine whether the interference correction range has been entered. Its value is determined by normalizing and fusing the overall similarity mutation index, grayscale distribution difference index, frame freeze index, and fixed noise energy mutation index. In specific implementation, it can be set and adjusted in the following ways: Value Range and Setting Logic: This threshold is typically set between 0.5 and 0.9. The principle is to ensure that, in sample sequences where camera correction has occurred, the fusion confidence of the starting frames of the vast majority of correction events (e.g., over 99%) exceeds this threshold; simultaneously, during normal monitoring periods without correction, the probability of confidence fluctuations in normal frames due to natural scene changes (e.g., cloud cover, object movement) exceeding this threshold should be extremely low (e.g., below 0.1%). This requires collecting a large amount of video from the target camera in typical application scenarios, statistically analyzing the confidence distribution of the correction event set and the normal scene set separately, and selecting a value that best distinguishes between the two as the initial threshold.
[0027] Adjustment principles: 1. Equipment dependence: Different models and brands of thermal imaging cameras have different degrees of correction severity and performance, so the optimal threshold may need to be calibrated specifically.
[0028] 2. Scene Adaptation: For applications with very stable scene thermal dynamics (such as monitoring an indoor server room), the threshold can be appropriately increased to increase the severity of judgment and reduce false triggers; for applications with complex scene thermal dynamics (such as monitoring an outdoor area where trees are swaying), the threshold can be appropriately decreased to ensure that all possible correction events can be captured and to avoid missed judgments.
[0029] 3. Online fine-tuning: During long-term operation, events with confidence levels exceeding the threshold can be recorded, and the threshold can be automatically fine-tuned based on whether there is a subsequent stable recovery process (i.e., it is determined to be a complete correction interference interval), thereby achieving adaptive optimization.
[0030] S2. Pre-build a device imaging noise fingerprint model, and select a background statistical region for each video frame based on the device imaging noise fingerprint model.
[0031] It should be noted that the device imaging noise fingerprint model aims to capture and quantify the inherent, scene-independent imaging defects of a specific thermal imaging camera, and can distinguish between image changes caused by real physical events (such as fires) and changes caused by the device's own noise.
[0032] In an optional embodiment, in S2, constructing the device imaging noise fingerprint model includes: During the preset fire-free period, for each pixel, the frequency of abnormally high values, the frequency of abnormal gray-level jumps, and the degree of gray-level consistency deviation with neighboring pixels are counted. It should be noted that the preset fire-free period refers to a period of time during which no real flames or smoke exist in the monitored scene, either during initialization or based on environmental confirmation (such as through manual confirmation or visible light image-assisted judgment). Video frames collected during this period are used for noise characteristic learning to ensure that the statistical results are not contaminated by the target signal.
[0033] It should be noted that, for each pixel, the frequency of abnormally high values refers to the proportion of times the pixel's grayscale value exceeds its long-term statistical mean (e.g., the average value within a sliding window) plus a certain multiple (e.g., 3-5 times) of the standard deviation (or exceeds an absolute high threshold) during periods without fire, out of the total number of statistical frames. This is used to identify overheated pixels that may be sensitive to high temperatures.
[0034] It should be noted that the frequency of abnormal grayscale jumps refers to the frequency at which the grayscale value of a pixel changes discontinuously and significantly between consecutive frames. Specifically, this can be achieved by calculating the absolute value of the difference between adjacent frames; when this value exceeds a dynamic threshold set based on the scene noise level, it is counted as a jump, and its occurrence frequency is recorded. This is used to identify flickering pixels with unstable responses.
[0035] It should be noted that the grayscale consistency deviation from neighboring pixels is used to measure the statistical consistency of the grayscale values of a pixel with its surrounding pixels (e.g., 8-neighborhood). In a uniform scene without fire, a normal pixel should have a high correlation with its neighborhood. This deviation can be obtained by calculating the multiple of the difference between the pixel's grayscale value and the mean of its neighborhood relative to the neighborhood's standard deviation, and counting the frequency with which its absolute value exceeds a threshold (e.g., 2.0). This is used to identify isolated bad pixels that are severely mismatched with their surroundings.
[0036] Based on the abnormal high frequency, the abnormal gray-level jump frequency, and the gray-level consistency deviation, the probability value of bad pixels for each pixel is calculated, and a bad pixel probability map is generated; wherein, the bad pixel probability map constitutes the imaging noise fingerprint model of the device.
[0037] It should be noted that calculating the probability value of a defective pixel for each pixel involves normalizing the three statistical indicators mentioned above (frequency of abnormally high values, frequency of abnormal jumps, and consistency deviation), and then performing a linear weighted fusion according to preset weights to obtain a probability value between 0 and 1. The closer this value is to 1, the higher the probability that the pixel is a defective pixel.
[0038] In an optional embodiment, in S2, a background statistical region is selected for each video frame based on the device imaging noise fingerprint model, as follows: Based on the bad pixel probability map, pixels whose bad pixel probability value exceeds the preset second threshold are excluded; Pixels within the moving region are excluded based on inter-frame difference information; It should be noted that excluding moving regions based on inter-frame difference information specifically refers to: calculating the absolute difference image between the current frame and the previous frame (or background model frame), and performing thresholding and morphological operations (such as erosion and dilation) on this difference image to generate a connected moving region mask. All pixels located within this mask are considered part of a moving object (such as a pedestrian or vehicle), and because their grayscale changes are not caused by a fixed scene or potential fire (early fires may be static), they are excluded from the background statistical candidates.
[0039] It should be noted that the preset second threshold is used to binarize the defective pixel probability values of pixels, distinguishing high-risk pixels. This threshold is typically set between 0.6 and 0.95. When setting this threshold, the defective pixel probability values of all pixels must be calculated within a preset fire-free period. Ideally, the probability values of the vast majority of normal pixels should be close to 0, while the probability values of a few inherently defective pixels (such as dead pixels or overheated pixels) should be close to 1. The initial value of this threshold should be set at a clear trough between the probability value distributions of these two types of pixels. For example, if statistics show that 95% of normal pixels have a probability value less than 0.2, while the probability values of all known physical defective pixels are greater than 0.8, then the threshold can be initially set to 0.7.
[0040] Adjustment principles: 1. Sensor health tolerance: If extremely sensitive to false alarms, the threshold can be increased (e.g., set to 0.9) to only block the most certain bad pixels, but some slightly abnormal pixels may remain. If extremely high image quality is required or the sensor is aging, the threshold can be decreased (e.g., set to 0.6) to more aggressively block any suspicious pixels and ensure the purity of the background statistical area, but some effective pixels may be lost.
[0041] 2. Dynamic Updates: The probability map of bad pixels itself becomes more accurate as the statistical period increases. In the early stages of system operation, a relatively conservative (high) threshold can be set, and as data accumulates, it can be adjusted to the optimal value based on a more stable probability distribution.
[0042] 3. Environmental reference: In extreme high and low temperature environments, the noise characteristics of the sensor may change. Consider making a small compensation to the threshold based on the camera temperature or ambient temperature.
[0043] Exclude locally highlighted areas and their morphologically dilated neighboring pixels; It should be noted that excluding locally highlighted areas and their morphologically dilated neighboring pixels aims to prevent the misuse of existing suspected high-temperature points or strong reflection sources as background. First, the image is segmented using a high grayscale threshold to obtain initial highlighted areas. Then, a morphological dilation operation (e.g., using a 3x3 or 5x5 structuring element) is used to expand the boundaries of these areas, ensuring complete exclusion of transition pixels that might be affected by hotspot edges. The dilation radius can be set empirically.
[0044] From the remaining pixels, the region with the smallest gray-level variance within a short time window is selected as the background statistical region.
[0045] It should be noted that selecting the region with the smallest gray-level variance within a short time window as the background statistical region is the final step in background selection. A short time window typically refers to a few recent consecutive frames, such as 10-30 frames. For each candidate pixel remaining after the aforementioned filtering steps, its gray-level variance within that time window is calculated. Finally, a continuous or discrete region with the smallest sum of variance (e.g., selecting several grids with the lowest variance after meshing the image) is selected as the most stable and reliable background statistical region for the current frame, used for subsequent gray-level mapping relationship estimation.
[0046] S3. When it is determined that the current frame is in the correction interference interval, the latest stable video frame before the start of the correction interference interval is determined as the reference frame. Based on the gray-scale statistical relationship between the background statistical region of the current frame and the reference frame, a constrained monotonic mapping function is constructed. The current frame is then adjusted by gray-scale mapping through the monotonic mapping function to generate a video frame with consistent gray-scale across the correction interference interval.
[0047] It should be noted that by bringing the grayscale statistical characteristics of the corrected frame back to the statistical baseline of the stable frame before correction, the continuity of time is restored at both the visual and algorithmic levels.
[0048] It should be noted that determining the latest stable video frame before the start of the interference correction interval as the reference frame is a key strategy. "Latest" ensures that the reference frame is temporally closest to the current scene (e.g., ambient temperature, lighting conditions), avoiding errors introduced by slow scene changes. "Stable" means that the frame itself is not within any identified interference correction interval, and the grayscale variance of its background statistical region is below a certain threshold, thus ensuring that its grayscale distribution is a reliable benchmark.
[0049] In an optional embodiment, in S3, a video frame with grayscale uniformity across the correction interference interval is generated, as follows: Extract a set of grayscale quantiles with the same percentile from the background statistical regions of the current frame and the reference frame respectively as anchor points; It's important to note that extracting a set of grayscale quantiles with the same percentile as anchor points is the core method for achieving black-box consistency. Common percentile choices include the 10th, 50th, and 90th percentiles. The 10th percentile (P10) represents dark features and is sensitive to shadows and low-temperature backgrounds; the 50th percentile (P50), the median, represents the overall brightness level; and the 90th percentile (P90) represents bright features and is sensitive to high-temperature areas and ambient heat sources. By comparing these statistical anchor point pairs (P10 of the current frame with P10 of the reference frame, and so on), the overall grayscale mapping shift, stretching, or compression relationship from dark to bright areas can be characterized without knowing the specific gain or shift parameters inside the camera.
[0050] Using the anchor point of the current frame as input and the anchor point of the reference frame as the target, a constrained piecewise linear monotonic function is constructed as a monotonic mapping function. It's important to note that constructing a constrained piecewise linear monotonic function means building a piecewise linear function connecting the extracted anchor point pairs as nodes. This function must be monotonically increasing to ensure that the mapping does not change the relative order of gray levels (i.e., hotter points remain hotter after mapping). The imposed constraints are crucial and typically include: 1. Slope Constraint: The slope of each segment of the linear function is constrained within a reasonable range (e.g., [0.5, 2.0]). A slope less than 1 indicates compression, and a slope greater than 1 indicates stretching. This constraint prevents extreme slopes caused by a few anomalous anchor points or noise, thus avoiding the amplification of small noise into significant spurious signals.
[0051] 2. Smoothness constraint: At the connection points of adjacent segments, the change of function value (or first derivative) should remain continuous or gradual to avoid drastic inflection points. This helps to suppress unnatural grayscale discontinuities or block effects in the mapped image.
[0052] Calculate the fitting residual of the monotonic mapping function and evaluate the reliability of the uniformization process by combining it with the coverage of the background statistical region. It should be noted that calculating the fitting residuals of the monotonic mapping function is crucial for evaluating the quality of the unification process. The residuals typically refer to the mean squared error (MSE) or mean absolute error (MAE) between the output values of all anchor point pairs after the mapping function transformation and the reference target value. Smaller residuals indicate a good fit of the statistical relationship. Evaluating the reliability of the unification process in conjunction with the coverage of the background statistical region means that, in addition to the fitting residuals, the proportion of background pixels used for fitting to the total number of pixels in the image (coverage) must also be considered. If the coverage is too low (e.g., below 20%), it indicates that the background area available for reliable statistics is too small. In this case, even if the fitting residuals are small, the risk of extrapolating the resulting mapping function to the entire image is high, and the reliability should be reduced. The reliability can be designed as a composite function of the fitting residuals and the coverage, for example, ,in To prevent division by zero for small constants.
[0053] When the confidence level is higher than the preset third threshold, the monotonic mapping function is applied to perform grayscale transformation on the current frame to generate a grayscale-consistent video frame.
[0054] It should be noted that the preset third threshold is used to determine whether to apply this mapping function. This threshold is calibrated experimentally: in a scenario where correction interference occurs but there is no real fire, the confidence values corresponding to a large number of successful and failed consistency attempts are statistically analyzed, and a cutoff value is selected accordingly. Only when the confidence value is higher than this threshold is the estimated mapping relationship considered reliable enough to be applied to the full-image grayscale transformation of the current frame, thereby generating a grayscale-consistent video frame. If the confidence value is insufficient, the process will proceed to the degradation detection strategy in step S5 to avoid the distortion and false alarm risks introduced by forced consistency.
[0055] It should be noted that the preset third threshold is a safety gate for determining whether to apply the estimated monotonic mapping function for grayscale adjustment. Its setting is crucial, directly affecting whether to perform high-quality frame consistency or safely switch to a degradation strategy. This threshold is typically set in the higher range of 0.7 to 0.95. Its setting is based on statistical analysis of the confidence evaluation metric. In a large number of samples that have successfully achieved effective consistency (demonstrated by the restoration of continuous statistical characteristics of the background region after correction), the calculated confidence should be concentrated in the higher range (e.g., >0.85). However, in scenarios where fitting fails, the background region is severely insufficient, or contaminated, the confidence is usually lower (e.g., <0.5). This threshold should be set in a safe zone between the low end of the confidence distribution of successful samples and the high end of the confidence distribution of failed samples.
[0056] Adjustment principles: 1. Trade-off between quality and safety: Increasing the threshold means stricter requirements for consistency quality, and the system is more inclined to adopt degradation strategies. This results in higher security but may sacrifice continuous optimization in some scenarios. Lowering the threshold will make the system more aggressive in trying to apply mapping functions, which may introduce risks in some edge cases.
[0057] 2. Dynamic weighting based on coverage and residuals: The credibility itself is composed of the fitted residuals and the coverage of the background statistical region. In actual deployment, the weights of these two components can be adjusted according to the application scenario. For example, in open scenarios, the background coverage is usually high, so more emphasis can be placed on the residuals; in complex occlusion scenarios, the coverage may often be low, in which case the stringent requirements for the residuals should be appropriately reduced, or the overall threshold should be adjusted accordingly.
[0058] 3. Linkage with the effectiveness of the degradation strategy: The setting of this threshold needs to be evaluated together with the actual effect of the degradation detection strategy. The goal is that above the threshold, the benefits (continuous improvement) brought by applying the mapping function significantly outweigh its potential risks; below the threshold, the reliability of the degradation strategy is sufficient to ensure basic fire detection performance.
[0059] S4. Based on the grayscale-uniformed video frames, extract fire features that are robust to grayscale mapping changes, and perform continuous evidence accumulation and fire determination across the correction interference interval.
[0060] It should be noted that reliable fire detection is achieved using image sequences that have undergone uniformization and are continuously comparable. Features insensitive to changes in global grayscale mapping are employed, and a mechanism for accumulation and judgment is designed to overcome interference and maintain the continuity of the evidence chain during in-camera correction.
[0061] In an optional embodiment, in S4, based on the grayscale-uniformed video frames, robust fire features to grayscale mapping changes are extracted as follows: The extracted fire features that are robust to grayscale mapping changes include: The difference between the high grayscale quantile and the median grayscale quantile within the candidate region; The difference between the average gray level of the candidate region and the average gray level of its surrounding annular background region; The trend of changes in the area, perimeter, or compactness of candidate hotspot connected regions over time.
[0062] It should be noted that robust fire features to grayscale mapping changes refer to fire features whose calculation relies on the relative ordering or local contrast of pixel grayscale values, rather than absolute grayscale values. Global linear or monotonic mapping transformations do not alter the nature of these features. The candidate regions are typically obtained by segmenting the image using a low initial temperature threshold and combining this with connected component analysis, representing potential high-temperature points or areas within the image.
[0063] It should be noted that the difference between the high-end gray-level quantile and the median gray-level quantile within the feature candidate region is robust. For example, calculate the difference (P90-P50) between the 90th percentile (P90) and the 50th percentile (P50, i.e., the median) within this region. This difference reflects the temperature difference between the hottest and typical parts of the region. Even if the entire image is brightened or darkened due to gain adjustment, as long as a significant temperature gradient exists within the fire area, this relative difference will remain stable, thus effectively indicating the presence of a local high-temperature core.
[0064] It should be noted that the difference between the average gray level of the feature candidate region and the average gray level of its surrounding annular background region utilizes local contextual contrast. Specifically, based on the bounding rectangle or minimum bounding circle of the candidate region, an annular region is formed by extending outwards by a certain pixel width (ensuring it does not overlap with the candidate region). The difference between the average gray level of the candidate region and the average gray level of the annular background region is calculated. This difference reflects the prominence of the hotspot relative to its immediate surrounding environment. Because this calculation is local and based on the difference, global gray-level shifts have little effect, and flames typically cause this difference to be significantly positive and persistent.
[0065] It should be noted that the changes in area, perimeter, or compactness of the connected regions of candidate hotspots over time are morphodynamic features. The increasing trends in area and perimeter are a direct visual manifestation of fire spread. Compactness is defined as... , is an index that measures how close a region's shape is to a circle (circular compactness is 1). In the early stages of a flame, it often exhibits irregular jumping motions, and its compactness may be low and fluctuating; during stable combustion or spread, it may show certain shape change patterns. The calculation of these morphological characteristics is entirely based on the spatial properties of the binary region and is independent of the absolute value of grayscale, thus being completely immune to changes in grayscale mapping.
[0066] In an optional embodiment, in S4, continuous evidence accumulation and fire situation determination are performed across the correction interference interval, as follows: Within the interference correction interval, pause the iteration of the background statistical model used for fire determination and disable evidence based on inter-frame differences; The fire evidence score is accumulated in the first time window before the start of the correction interference interval and the second time window after the end of the correction interference interval, respectively. In an optional embodiment, the fire evidence scores are accumulated as follows: For the video frame sequence within the first time window, the first evidence score is accumulated using a sliding window method; For the video frame sequence within the second time window, the second evidence score is accumulated using a sliding window method.
[0067] It should be noted that the accumulated fire evidence scores within the first time window before the start of the correction interference interval and the second time window after its end are the core of the continuous accumulation across intervals. The evidence accumulated in the first time window (e.g., 5-10 seconds before the interference begins) represents the initial state of the fire before the interference occurs. The evidence accumulated in the second time window (e.g., 5-10 seconds after the interference ends) represents the recovered state of the fire after the interference ends. This separate accumulation ensures that the evidence in both stages remains uncontaminated, providing clean input for subsequent fusion.
[0068] It should be noted that accumulating evidence scores using a sliding window method means that, within each time window, for each frame, the robust features extracted above are converted into a single-frame evidence score between 0 and 1 (for example, by mapping feature differences to probabilities using a sigmoid function). Then, within a sliding time window (e.g., 3 seconds), these single-frame scores are integrated (summed) or the proportion of their consistently high values is calculated to obtain a more stable window evidence score that is resistant to transient noise (i.e., the first evidence score or the second evidence score).
[0069] It should be noted that pausing the iteration of the background statistical model used for fire situation determination means stopping the use of current frame information to update the parameters used for adaptive background modeling (such as Gaussian mixture model) within the interference correction interval. This is because the grayscale values of the frame are distorted at this time, and using them to update the background model would pollute the model, leading to continuous false alarms or missed alarms after the interference ends.
[0070] It's important to note that evidence based on inter-frame differences is disabled because during abrupt changes in grayscale mapping, the difference between consecutive frames primarily reflects the global changes introduced by the correction, rather than the actual motion or intensity changes of objects within the scene. Using this difference would directly lead to misinterpreting the correction as a significant change event, resulting in serious false alarms.
[0071] A fusion judgment is made based on the accumulated fire evidence scores within the first and second time windows.
[0072] In an optional embodiment, the fire evidence scores are accumulated as follows: The first evidence score and the second evidence score are weighted and summed, or the difference in the growth of the second evidence score relative to the first evidence score is compared, to obtain the final basis for determining the fire situation.
[0073] It should be noted that the two main methods for fusion determination have their applicable scenarios: 1. Weighted summation: Add the first evidence score (S1) and the second evidence score (S2) according to their weights (e.g.: This applies to situations where a comprehensive assessment of the fire's duration is required, and the weighting can reflect whether more importance is placed on evidence from before or after the interference.
[0074] 2. Compare the growth difference: Calculate .if A value greater than a positive threshold indicates that despite the correction interference, the fire evidence not only recovers after the interference ends but is also stronger than before, which closely matches the expected development of a real fire (fire spread). This approach is insensitive to slowly changing heat sources (such as continuously operating machinery) because they typically do not cause... A significant positive growth. Finally, the total score or difference obtained based on the selected fusion method is compared with a final alarm threshold to determine whether to trigger a fire alarm.
[0075] S5. When the estimation confidence of the monotonic mapping function is insufficient, a downgrade detection strategy is enabled. The downgrade detection strategy includes: skipping the grayscale mapping adjustment of the current frame and directly determining the fire situation based on the fire features and fusing the separation evidence before and after the correction interference interval.
[0076] It should be noted that this step ensures that monitoring functions can still be maintained in a conservative but effective manner when the core consistency processing cannot be executed securely due to insufficient reliability. This avoids the risk of false alarms caused by forcibly applying incorrect mappings or the risk of missed alarms caused by a complete interruption of analysis.
[0077] It should be noted that insufficient reliability of the monotonic mapping function estimate is the decision condition for triggering the degradation strategy. This typically occurs in scenarios such as extremely low coverage of the background statistical region, excessively large fitting residuals, or severe bad spot contamination, causing the reliability calculated in step S3 to fall below a preset third threshold. This threshold is calibrated through extensive experiments to identify situations where the mapping relationship estimate is extremely unreliable and its application risk outweighs the benefits.
[0078] It should be noted that the core of the degradation strategy, skipping the grayscale mapping adjustment of the current frame, means that for frames within the correction interference range and with insufficient confidence, pixel-level grayscale transformation is abandoned. These frames will not participate in constructing a visually continuous video stream, but will be marked as low-quality frames or frames that cannot be directly compared.
[0079] It should be noted that the core analysis logic in the downgrade mode is to directly determine the fire situation based on the fire characteristics and to fuse the separate evidence before and after the correction interference interval. In this mode, it relies entirely on robust features defined in step S4 that are insensitive to changes in grayscale mapping, such as quantile difference, local contrast, and morphological trends. For frames within the correction interference interval, since mapping adjustments are skipped, their feature calculations may be based on distorted original grayscale values, but robust design minimizes the impact. Evidence fusion strictly follows the cross-interval accumulation principle described in step S4: scores are strictly accumulated separately in two independent evidence windows before and after the correction interference, and then fusion judgments are performed, such as weighted summation or comparison of differences. This ensures that even without reliable evidence in the intermediate interval (correction interference period), a coherent judgment can be made based on the event precursors and consequences, effectively preventing the complete break in the chain of evidence.
[0080] In an optional embodiment, in S5, the degradation detection strategy further includes: When the confidence level of the estimate is insufficient, the location information of the suspected fire area identified before the correction interference interval is retained, and the suspected fire area is re-verified after the correction interference interval ends.
[0081] It should be noted that the optional strategy of retaining and re-verifying the location information of suspected fire areas is an enhanced protection mechanism for early-stage small fires. Specifically, before the correction interference begins, if the fire characteristics (such as local contrast) of a certain area have shown some significance but have not yet reached the alarm threshold, the system caches the geographical coordinates (image pixel coordinates), circumscribed rectangle, or contour information of that area. After exiting the correction interference interval and resuming normal processing, in subsequent frames, the cached location area will be scanned and analyzed first to check whether the characteristics of that area persist or develop. This is equivalent to providing a memory and rapid review channel for signs that appeared before the interference occurred, significantly improving the early detection probability of small fires that may continue to develop during the interference period.
[0082] S6. Output the fire situation determination result and associated image quality markers.
[0083] The output image quality markers include at least one of the following: correction interference interval markers, monotonic mapping function estimation confidence, and the proportion of high-risk pixels in the bad pixel probability map.
[0084] It should be noted that sending alarm information containing the image quality markers to the remote monitoring platform means that the alarm information is a structured data packet.
[0085] It contains at least: 1. Key event data: Device ID, alarm time, fire area coordinates / image screenshot; 2. Process quality information: the state of correction interference when the alarm is triggered, the reliability of the mapping function used, the sensor health status (proportion of bad pixels), etc. 3. Evidence summary: The cumulative evidence score curve or key value across intervals.
[0086] This information is reported together, allowing remote monitoring personnel not only to know what happened, but also to understand the data quality conditions under which the system made its judgments, thus enabling them to conduct more evidence-based reviews and decisions.
[0087] When a fire is detected, an alarm message containing the image quality marker is sent to the remote monitoring platform, and a linkage control signal is triggered.
[0088] It should be noted that triggering a linkage control signal refers to the system sending standard control commands to pre-configured security equipment locally or via a network. For example, it may trigger the sounding and flashing of on-site audible and visual alarms, activate the emergency broadcast system to play evacuation instructions, or send a pre-action signal to an automatic fire extinguishing system (such as a gas extinguishing controller) via dry contact signals or network protocols. This linkage achieves a closed loop from intelligent sensing to automatic response, shortening emergency response time.
[0089] In summary, this patent application solves the interference problem introduced by equipment self-calibration in thermal imaging fire monitoring through a series of synergistic technical means, achieving continuous monitoring, robust judgment, system security, and interpretability of output, and significantly improving the accuracy and reliability of early fire warning.
[0090] Furthermore, any content not described in detail in this specification is existing technology known to those skilled in the art.
[0091] In the embodiments provided by this invention, it should be understood that the disclosed system or method can be implemented in other ways. For example, the embodiments of the invention described above are merely illustrative; for instance, the division of modules is only a logical functional division, and there may be other division methods in actual implementation.
[0092] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0093] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated module can be implemented in hardware or in the form of hardware plus software functional modules.
[0094] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the basic characteristics of the present invention.
[0095] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. An intelligent remote monitoring method for fire detection based on image processing, characterized in that, Includes the following steps: S1. Acquire consecutive image frames of thermal imaging video, construct a thermal imaging video sequence, and identify the correction interference interval where the image grayscale mapping changes abruptly based on the thermal imaging video sequence. S2. Pre-build a device imaging noise fingerprint model, and select a background statistical region for each video frame based on the device imaging noise fingerprint model; S3. When it is determined that the current frame is in the correction interference interval, the latest stable video frame before the start of the correction interference interval is determined as the reference frame; based on the gray-scale statistical relationship between the background statistical region of the current frame and the reference frame, a constrained monotonic mapping function is constructed, and the gray-scale mapping of the current frame is adjusted through the monotonic mapping function to generate a video frame with consistent gray-scale across the correction interference interval. S4. Based on the grayscale-uniformed video frames, extract fire features that are robust to grayscale mapping changes, and perform continuous evidence accumulation and fire determination across the correction interference interval. S5. When the estimation confidence of the monotonic mapping function is insufficient, a downgrade detection strategy is enabled. The downgrade detection strategy includes: skipping the grayscale mapping adjustment of the current frame and directly determining the fire situation based on the fire features and fusing the separation evidence before and after the correction interference interval. S6. Output the fire situation determination result and associated image quality markers.
2. The intelligent fire monitoring and remote control method based on image processing as described in claim 1, characterized in that: In S1, the correction interference interval for sudden changes in image grayscale mapping is identified based on the thermal imaging video sequence, as follows: Calculate the overall similarity abruptness index between the current frame image and the previous frame image; Calculate the grayscale distribution difference index between the current frame image and the previous frame image; Detect frame freeze indicators in a video sequence; Detect abrupt changes in fixed-pattern noise energy in an image and generate a fixed-noise energy abrupt change index; Based on the overall similarity mutation index, the grayscale distribution difference index, the frame freeze index, and the fixed noise energy mutation index, a correction interference confidence level is generated. When the confidence level of the correction interference exceeds a preset first threshold, it is determined that the frame has entered the correction interference interval.
3. The intelligent fire monitoring and remote control method based on image processing as described in claim 1, characterized in that: In S2, constructing the device imaging noise fingerprint model includes: During the preset fire-free period, for each pixel, the frequency of abnormally high values, the frequency of abnormal gray-level jumps, and the degree of gray-level consistency deviation with neighboring pixels are counted. Based on the abnormal high frequency, the abnormal gray-level jump frequency, and the gray-level consistency deviation, the probability value of bad pixels for each pixel is calculated, and a bad pixel probability map is generated; wherein, the bad pixel probability map constitutes the imaging noise fingerprint model of the device.
4. The intelligent fire monitoring and remote control method based on image processing as described in claim 3, characterized in that: In S2, a background statistical region is selected for each video frame based on the device imaging noise fingerprint model, as follows: Based on the bad pixel probability map, pixels whose bad pixel probability value exceeds the preset second threshold are excluded; Pixels within the moving region are excluded based on inter-frame difference information; Exclude locally highlighted areas and their morphologically dilated neighboring pixels; From the remaining pixels, the region with the smallest gray-level variance within a short time window is selected as the background statistical region.
5. The intelligent fire monitoring and remote control method based on image processing as described in claim 1, characterized in that: In S3, video frames with uniform grayscale across the correction interference interval are generated as follows: Extract a set of grayscale quantiles with the same percentile from the background statistical regions of the current frame and the reference frame respectively as anchor points; Using the anchor point of the current frame as input and the anchor point of the reference frame as the target, a constrained piecewise linear monotonic function is constructed as a monotonic mapping function. Calculate the fitting residual of the monotonic mapping function and evaluate the reliability of the uniformization process by combining it with the coverage of the background statistical region. When the confidence level is higher than the preset third threshold, the monotonic mapping function is applied to perform grayscale transformation on the current frame to generate a grayscale-consistent video frame.
6. The intelligent fire monitoring and remote control method based on image processing as described in claim 1, characterized in that: In S4, based on the grayscale-uniformed video frames, robust fire features to grayscale mapping changes are extracted as follows: The extracted fire features that are robust to grayscale mapping changes include: The difference between the high grayscale quantile and the median grayscale quantile within the candidate region; The difference between the average gray level of the candidate region and the average gray level of its surrounding annular background region; The trend of changes in the area, perimeter, or compactness of candidate hotspot connected regions over time.
7. The intelligent fire monitoring and remote control method based on image processing as described in claim 6, characterized in that: In S4, continuous evidence accumulation and fire situation determination are performed across the correction interference interval, as follows: Within the interference correction interval, pause the iteration of the background statistical model used for fire determination and disable evidence based on inter-frame differences; The fire evidence score is accumulated in the first time window before the start of the correction interference interval and the second time window after the end of the correction interference interval, respectively. A fusion judgment is made based on the accumulated fire evidence scores within the first and second time windows.
8. The intelligent fire monitoring and remote control method based on image processing as described in claim 7, characterized in that: The scores for fire evidence are calculated separately as follows: For the video frame sequence within the first time window, the first evidence score is accumulated using a sliding window method; For the video frame sequence within the second time window, the second evidence score is accumulated using a sliding window method.
9. The intelligent fire monitoring and remote control method based on image processing as described in claim 8, characterized in that: The scores for fire evidence are calculated separately as follows: The first evidence score and the second evidence score are weighted and summed, or the difference in the growth of the second evidence score relative to the first evidence score is compared, to obtain the final basis for determining the fire situation.
10. The intelligent fire monitoring and remote control method based on image processing as described in claim 1, characterized in that: In S5, the degradation detection strategy also includes: When the confidence level of the estimate is insufficient, the location information of the suspected fire area identified before the correction interference interval is retained, and the suspected fire area is re-verified after the correction interference interval ends.