Intelligent fire detection method based on artificial intelligence video analysis

By using an AI-based video analysis-based intelligent fire detection method, the problem of difficulty in identifying initial flames and fine smoke in existing technologies has been solved, achieving high-confidence fire identification and real-time alarm, and is suitable for fire early warning in complex scenarios.

CN121033734AActive Publication Date: 2025-11-28SHANGHAI ZHISHENG INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202511563790.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2025-11-28
Estimated Expiration
2045-10-30

AI Technical Summary

Technical Problem

Existing video-based fire detection methods struggle to identify weak flames and fine smoke in the initial fire stage, and suffer from high false positive and false negative rates in complex environments, making it difficult to meet the high reliability requirements of industrial applications.

Method used

An intelligent fire detection method based on artificial intelligence video analysis is adopted. By collecting video surveillance streams, spatial color and morphological features of the images are extracted, temporal difference analysis is performed, multi-scale temporal feature maps are constructed, and temporal feature extraction is performed using a dual-channel neural network model. Combined with boundary stability analysis and a multi-modal fusion discrimination model, the fire judgment result is output.

Benefits of technology

It achieves high-confidence determination of fire areas, reduces false alarm and false alarm rates, and can output the fire initiation location and spread trend in real time, providing a visual decision-making basis for fire dispatch. It is suitable for complex scenarios such as smart cities and industrial parks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121033734A_ABST
    Figure CN121033734A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent fire detection method based on artificial intelligence video analysis, and particularly relates to the technical field of computer vision. The method comprises the following steps: acquiring a video monitoring stream, extracting spatial color features and morphological features of an image, constructing a multi-scale time sequence feature graph group, performing joint modeling on color disturbance and structural disturbance by using a dual-channel attention neural network, and outputting a suspected fire area probability graph; candidate fire source areas are screened in combination with boundary disturbance consistency analysis and red channel high-frequency fluctuation detection; the texture stability, the disturbance directivity and the historical smoke evolution characteristics are integrated through a multi-modal fusion judgment model, the fire state is finally judged, an alarm result is output, meanwhile, the system can mark the fire position and the development trend path, and visual tracking is achieved. The method has the advantages of high robustness, high sensitivity and low false alarm rate, and is suitable for intelligent early warning of early fire in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision, in particular to a fire intelligent detection method based on artificial intelligence video analysis. BACKGROUND

[0002] With the improvement of automation and unmanned level in closed or semi-closed environments such as urban underground space, warehouse logistics center, and large-scale unmanned logistics park, the traditional fire detection technology relying on smoke and temperature sensors gradually exposes many limitations in practical application. For example, in the unattended closed scene, smoke is difficult to diffuse to the sensor sampling point in time, and false alarms are prone to occur in high-temperature environments (such as industrial drying and logistics oven), causing unnecessary production stoppage and resource waste.

[0003] In recent years, the rapid development of artificial intelligence image recognition technology makes it possible to identify fire based on video. However, most of the existing fire identification methods based on video only rely on convolutional neural network models to classify image features, lack modeling of the dynamic fire evolution process in the time dimension, and it is especially difficult to identify weak light and subtle smoke in the initial flame stage, which is the key early warning opportunity before the spread of fire. In addition, there are complex factors such as strong light interference, mechanical motion artifacts, and hot gas disturbance in the video, which makes the false detection rate of traditional image algorithms high and the missed detection rate large in early identification, and it is difficult to meet the high reliability requirements of industrial applications. Especially in special scenes such as automatic stereoscopic warehouse, logistics conveyor fire source monitoring, and unmanned power distribution room, the existing technology has almost no effective solution. SUMMARY

[0004] The purpose of the present application is to provide a fire intelligent detection method based on artificial intelligence video analysis to solve the problems in the background art.

[0005] In order to achieve the above purpose, the present application provides the following technical scheme: a fire intelligent detection method based on artificial intelligence video analysis, comprising: S100: collecting video monitoring flow of a target area, and extracting spatial color feature vector C and morphological feature vector M of each frame of image; S200: performing time sequence difference analysis on consecutive frames of images, obtaining change amplitude vectors AC and AM based on C and M, and establishing an initial dynamic feature set D; S300: constructing a multi-scale time sequence feature map group T, which is composed of three-dimensional tensors formed by D of consecutive N frames of images under different time windows; S400: using a double-channel neural network model based on attention mechanism to perform time sequence feature extraction on T, and outputting a suspected fire area probability map P; S500: Perform boundary stability analysis on the suspected region in P to determine whether there is a boundary sequence S with flame-shaped perturbation characteristics, and calculate its perturbation consistency coefficient R. S600: If R exceeds the set threshold, further combine the high-frequency fluctuation characteristics of the red channel in C and ΔC to generate candidate fire source marking regions F; S700: Input F into the fire multimodal fusion discrimination model, fuse the texture stability, perturbation directionality and historical smoke evolution model of region F under different lighting channels, and output the final fire judgment result L; S800: If L determines that a fire has occurred, an alarm signal is sent and the fire's starting position and development path are marked in the video frame.

[0006] Preferably, S200 includes: S201: Perform inter-frame registration processing on multiple consecutive frames of images within a preset time window; S202: Based on the registered image frame sequence, perform frame-by-frame difference operation on the spatial color feature vector C and morphological feature vector M of each frame image to calculate the color change amplitude ΔC and morphological change amplitude ΔM between adjacent frames. The difference method includes Euclidean distance or structural similarity calculation. S203: Perform filtering and smoothing on ΔC and ΔM; S204: Combine the smoothed ΔC and ΔM to construct a dynamic feature set D, which is used to describe the evolution characteristics of suspected flame or smoke regions in a video frame sequence.

[0007] Preferably, S300 includes: S301: Set multiple time window scale groups n is the total number of windows, and each time window corresponds to a frame interval, which is used to extract image sequence features of different time lengths from the dynamic feature set D. S302: For each time window, extract the feature map sequence of consecutive frames from the dynamic feature set D, and stack the sequence in time order to form a primary three-dimensional tensor Ti with size Wi×H×W, where H and W are the image height and width; S303: Perform temporal normalization on each primary 3D tensor Ti to generate a temporally normalized tensor Ti′; S304: Combine the time-normalized tensors Ti′ at all scales to form a multi-scale time-series feature map group T, which is used to characterize the dynamic evolution characteristics of fire at different time granularities.

[0008] Preferably, S400 includes: S401: Construct a dual-channel neural network model with a parallel structure, which receives color dynamic information and morphological dynamic information from the multi-scale temporal feature map group T respectively, and independently extracts the corresponding temporal spatial feature tensors; S402: Within each channel, calculate the keyframe weights in the time dimension of the input 3D tensor. S403: Perform channel fusion processing on the feature tensors output from the two channels to generate a fused multidimensional fire feature map; S404: The fused feature map is input into the spatial saliency decoding network, and convolutional and deconvolutional layers are used to gradually restore the suspected fire area probability map P with the same size as the original image. The value of each pixel in P represents the confidence probability that the location is a fire area.

[0009] Preferably, S500 includes: S501: Binarize the probability map P of suspected fire areas in multiple consecutive frames, extract the boundaries of areas with confidence scores higher than the first threshold, and extract the boundary contour set B for each frame. t ; S502: On the time series, the boundary contour set B t Inter-frame matching is performed, and boundary correspondence is constructed based on the geometric centroid position of the contour, the similarity of the boundary shape, and the area change ratio to form a candidate boundary sequence S; S503: Quantify the perturbation behavior of each boundary in the boundary sequence S and construct a perturbation feature vector set, including boundary deformation rate, centroid drift vector and local edge irregularity index; S504: Calculate the perturbation consistency coefficient R based on the perturbation feature vector group. R is the normalized result of the standard deviation of the perturbation features during the boundary change over time. If R is less than the second threshold, the boundary sequence is determined to have flame perturbation consistency.

[0010] Preferably, S600 includes: S601: Provided that the perturbation consistency coefficient R exceeds a set threshold, extract the spatial color features C of each frame from the image region corresponding to the boundary sequence S, and focus on extracting the red channel values. Pixel distribution; S602: Perform frequency domain analysis on the time-series value sequence of the red channel in consecutive frame images, extract the high-frequency components of the red channel, and construct the red high-frequency energy spectrum E; S603: Within the region corresponding to the boundary sequence, analyze the spatial distribution of the high-frequency energy spectrum E, identify the high-frequency fluctuation clustering region, and screen out the red high-frequency anomaly region R_high; S604: Perform spatial overlap calculation on the boundary sequence S and R_high, and extract the intersection region as the candidate fire source marker region F.

[0011] Preferably, S700 includes: S701: Extract texture stability index for candidate fire source marking region F under multiple illumination channels, and calculate its local texture consistency score in the original image, grayscale image and gamma-corrected image respectively; S702: Based on the contour deformation trajectory of region F in consecutive frames, construct a perturbation directionality vector set, obtain the main perturbation direction through principal component analysis, and calculate the direction perturbation deviation value. S703: Call the historical smoke evolution model to perform time-reverse feature comparison on region F, assess whether there is a continuous smoke diffusion process in its surrounding areas, and output a matching score; S704: The texture stability, perturbation directionality and historical smoke matching score are taken as inputs and fed into the fire multimodal fusion discrimination model for joint decision-making, and the final fire judgment result L is output, where L is a binary result or confidence score.

[0012] Preferably, S800 includes: S801: When the fire judgment result L is in the established state, generate the alarm time, target area location and alarm level, and send an alarm signal; S802: Starting from the time point when the fire was determined, trace back through consecutive video frames, extract the geometric center position of the candidate fire source marking region F in each frame image, and construct the fire starting path trajectory P_start; S803: In consecutive frames after the fire occurs, dynamically track the area expansion, boundary movement trend and center drift direction of region F, and calculate the fire spread path trajectory P_expand; S804: P_start and P_expand are concatenated to generate the fire development trend path P_total, and the fire source origin, movement direction and development trend are marked with trajectory lines in the alarm video frame.

[0013] The technical effects and advantages provided by the present invention in the above technical solution are as follows: 1. This invention provides an intelligent fire detection method based on artificial intelligence video analysis, which effectively overcomes the limitations of traditional methods relying on smoke sensors or static image recognition algorithms. It integrates multi-scale temporal feature extraction, boundary perturbation modeling, color frequency domain analysis, and a multimodal intelligent discrimination mechanism, possessing strong early fire source identification capabilities. By introducing a dual-channel neural network structure and attention mechanism, this invention can achieve high-confidence determination of fire areas in complex backgrounds, high-interference, or weak flame stages, significantly reducing false alarm and false negative rates, and significantly improving the system's adaptability and reliability in unmanned, high-risk scenarios.

[0014] 2. This invention constructs a global-local combined fire behavior modeling mechanism by integrating texture stability analysis, disturbance direction modeling, and historical smoke evolution trajectory backtracking. This mechanism not only intelligently determines the fire's establishment status but also outputs the fire's location, spread trend, and complete trajectory clues in real time, providing visualized decision-making support for fire dispatch, emergency response, and accident backtracking. The overall solution boasts advantages such as high intelligence, flexible deployment, and strong compatibility, making it particularly suitable for complex scenarios with high fire early warning accuracy requirements, such as smart cities, industrial parks, underground facilities, and logistics warehouses. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0016] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] For examples, please refer to Figure 1 As shown in this embodiment, the intelligent fire detection method based on artificial intelligence video analysis includes: S100: Acquires video surveillance stream of the target area and extracts the spatial color feature vector C and morphological feature vector M of each frame image; S200: Perform temporal difference analysis on consecutive frame images to obtain the change magnitude vectors ΔC and ΔM based on C and M, and establish an initial dynamic feature set D; S300: Construct a multi-scale temporal feature map group T, wherein T is composed of a three-dimensional tensor formed by the D of N consecutive frames of images under different time windows; S400: Use a dual-channel neural network model based on attention mechanism to extract temporal features from T and output a probability map P of suspected fire areas; S500: Perform boundary stability analysis on the suspected region in P to determine whether there is a boundary sequence S with flame-shaped perturbation characteristics, and calculate its perturbation consistency coefficient R. S600: If R exceeds the set threshold, further combine the high-frequency fluctuation characteristics of the red channel in C and ΔC to generate candidate fire source marking regions F; S700: Input F into the fire multimodal fusion discrimination model, fuse the texture stability, perturbation directionality and historical smoke evolution model of region F under different lighting channels, and output the final fire judgment result L; S800: If L determines that a fire has occurred, an alarm signal is sent and the fire's starting position and development path are marked in the video frame.

[0019] In this embodiment, step S100 aims to extract basic image feature information from the video source for subsequent fire detection analysis, serving as input data for the entire intelligent detection process. Specifically, it includes the following steps: First, real-time or stored video surveillance streams are acquired from the designated target monitoring area. These video surveillance streams are continuous frame image sequences, and the frame rate range can be configured according to the actual application scenario, preferably between 15 and 30 frames per second, to ensure the integrity and real-time nature of the timing information.

[0020] Next, each frame of the video surveillance stream is preprocessed. The preprocessing includes steps such as image normalization, noise suppression, brightness balancing, and resolution unification to eliminate interference with image quality caused by different device acquisition conditions and ensure the accuracy of subsequent feature extraction.

[0021] After preprocessing, two types of core feature information are extracted from each frame of the image: Spatial color feature vector C: C is a set of vectors representing the color distribution of each pixel in the image, preferably represented by the HSV (Hue-Saturation-Value) color space, where H represents the hue component, S represents the saturation component, and V represents the lightness component. This color feature can effectively reflect the typical red-orange fluctuation characteristics of the flame area in the image and has good adaptability to different lighting conditions. Optionally, to enhance the stability of color recognition, C may also include the mean and higher-order statistical features of the red channel (R) in the RGB space, such as skewness and kurtosis.

[0022] Morphological feature vector M: M is a set of features describing changes in the structural morphology of an image. Preferably, edge contour information of the image is obtained through edge detection operators (such as Canny or Sobel), and morphological features are calculated by combining the image gradient direction and texture density. This feature is used to assist in identifying the flickering boundaries and irregular disturbances of flames, especially in low-light or smoke-obscured scenes, providing a basis for structural stability analysis.

[0023] In one embodiment of the present invention, step S200 is used to extract spatial color feature changes and morphological structure changes from consecutive frame images, establish an initial dynamic feature set D reflecting the dynamics of fire development, and provide data support for subsequent time-series modeling and intelligent discrimination. The step specifically includes the following sub-steps: S201: Perform inter-frame registration processing on multiple consecutive frames of images within a preset time window.

[0024] Preferably, a dense optical flow method is used to estimate the motion vectors of all pixels in adjacent image frames, and subsequent frames are resampled in reverse to align them with the reference frame. Taking the Farneback algorithm as an example, the core idea of ​​the dense optical flow method is to estimate the pixel movement trend by expanding the polynomial within the image window, thereby generating an optical flow map between consecutive frames.

[0025] In the specific implementation, let the t-th frame be the reference frame and the (t+1)-th frame be the frame to be registered. Then, for each pixel (x, y), calculate the offset vector (dx, dy) between the two frames. Based on this vector, perform an affine transformation or perspective transformation on the (t+1)-th frame to complete the image alignment process.

[0026] S202: Based on the registered image frame sequence, perform frame-by-frame difference operation on the spatial color feature vector C and morphological feature vector M of each frame image to calculate the color change amplitude ΔC and morphological change amplitude ΔM between adjacent frames. The difference method includes Euclidean distance or structural similarity calculation.

[0027] After image frame registration is completed, the spatial color feature vector C and morphological feature vector M are extracted for each frame image, and the difference calculation is performed on the feature vectors between adjacent frames to obtain the color change amplitude ΔC and the morphological change amplitude ΔM.

[0028] Specifically, the color change amplitude ΔC is calculated as follows: Let the spatial color features of the t-th frame be... The image in frame t+1 is ,but The distance function is preferably a weighted Euclidean distance, which has the following form: , Where H, S, and V are the components of the HSV color space, respectively. The weighting coefficients set for the experience are preferably 0.4, 0.3, and 0.3.

[0029] Similarly, the calculation of the morphological change amplitude ΔM is based on the structural similarity (SSIM) of the image edge gradient: Let the edge map of the t-th frame be... The (t+1)th frame is ,but , SSIM is a structural similarity index with a value ranging from 0 to 1; a smaller value indicates a greater difference. This processing method ensures the extraction of key dynamic information about how flames and smoke change over time.

[0030] S203: To suppress feature errors caused by short-term noise, lighting jitter, or slight background disturbances, time smoothing processing is required for ΔC and ΔM.

[0031] Preferably, a weighted average method based on a sliding window is used. The specific method is as follows: Let the sliding window length be k, then the smoothed ΔC at time t is: , in , These are weighting coefficients; The weighting coefficients can be set according to a time-increasing function, preferably using an exponential decay form, such as: , where α∈(0,1), and preferably takes the value of 0.2.

[0032] Correspondingly, the same moving weighted average process is also performed on ΔM to generate a smoothed ΔM′.

[0033] S204: Jointly encode the smoothed color change vector ΔC′ and the shape change vector ΔM′ to form the initial dynamic feature set D.

[0034] In the specific implementation: the dynamic feature set D is a sequence of two-dimensional feature maps, where each frame of D... t This is a set of numerical matrices representing the distribution of ΔC′ and ΔM′ in the image space; ΔC′ and ΔM′ can be encoded by proportional superposition using a pixel-level fusion method. For example: λ is the fusion weighting factor, which is preferably set to 0.6.

[0035] In one embodiment of the present invention, step S300, in order to enhance the dynamic modeling capability of the fire occurrence process in the time dimension, employs a multi-scale modeling approach to temporally stack the aforementioned dynamic feature set D, constructing a multi-scale temporal feature map group T. The process includes the following steps: S301: Specifically, a time window scale group {W1, W2, ..., Wn} is defined, where n is the total number of windows, and each Wi represents the length of a time window in frames. Preferably, W1 = 3 (short-time window), W2 = 5 (medium-time window), and W3 = 7 (long-time window); the window corresponding to each time scale Wᵢ will be used to extract the image feature sequence of consecutive Wᵢ frames from the dynamic feature set D for subsequent construction of the three-dimensional tensor.

[0036] S302: For each time window Wi, obtain the feature map sequence of consecutive Wi frames preceding the current time from the dynamic feature set D, represented as follows: Each frame D in the sequence is a two-dimensional matrix, representing the combined result of the smoothed color change amplitude ΔC and the shape change amplitude ΔM. Stacking this Wi-frame sequence in chronological order on the first dimension of a tensor forms a three-dimensional tensor. Its shape is: Where H represents the image height and W represents the image width. Let R be the number of time frames and R be the set of real numbers. This three-dimensional tensor reflects the dynamic changes at this time scale and is an important foundation for time series modeling.

[0037] S303: Because the degree of change of image frames varies in different time periods, the tensor obtained by direct stacking may have an imbalance in time components. Therefore, the tensor needs to be... Temporal normalization and attention enhancement processes are performed to improve the saliency of key temporal features. This process consists of two parts: Inter-frame difference normalization processing: for tensors Any two adjacent frames and Calculate the difference Then for all Normalization ensures that the difference ranges between 0 and 1. Min-max normalization can be used as a normalization method. Where min and max are respectively The minimum and maximum values ​​in.

[0038] Enhanced Channel Attention Mechanism: To highlight keyframe information in the temporal dimension of the tensor, a channel attention mechanism, such as the Squeeze-and-Excitation (SE) module, is introduced.

[0039] The implementation steps are as follows: For tensors Perform global average pooling in the spatial dimension to obtain... A 3D vector; this vector is compressed and activated through a two-layer fully connected network, outputting a weighted coefficient vector. Use this coefficient to Frame-by-frame weighting is performed in the time dimension, that is, each frame is multiplied by the corresponding attention weight. The resulting output tensor This will highlight important frames in the time sequence and suppress background noise interference.

[0040] S304: Tensor normalized across all scales Joint encoding is performed to form the final multi-scale temporal feature map group T.

[0041] The specific method is as follows: for all scales By concatenating the data in either the "time dimension" or the "channel dimension" (e.g., using a concat operation), a tensor input of uniform shape is obtained, which is then used in subsequent deep learning network models, such as 3D convolutional neural networks (3D-CNN) or convolutional long short-term memory networks (ConvLSTM). The final multi-scale feature map set T can comprehensively reflect the dynamic evolution characteristics of the flame at different time lengths, including: the frequency of flame outline changes; the directionality of smoke diffusion; and the periodicity of brightness and darkness.

[0042] In this embodiment, step S400 is used to perform deep learning processing on the aforementioned constructed multi-scale temporal feature map group T. By introducing an attention mechanism and a dual-channel structure, high-confidence identification of fire areas is achieved, and finally, a probability map P of suspected fire areas is output. The specific steps are as follows: S401: This step constructs a dual-channel convolutional neural network model with a parallel structure, which includes: Color dynamic channel: used to receive and process inputs containing color change features (such as ΔC, HSV channel changes, etc.) in feature map group T; Morphological dynamic channel: used to process feature inputs containing morphological changes (such as edge perturbation, texture jumps, etc.) (such as ΔM, gradient map, etc.).

[0043] Each channel contains a feature extraction subnetwork composed of stacked 3D convolutional layers (3D-CNN). 3D convolution can simultaneously extract spatial (image structure) and temporal (inter-frame evolution) features, making it an ideal choice for analyzing dynamic events such as fires.

[0044] Taking color channels as an example, the model structure is as follows: Input dimensions: Let represent the number of time frames as n, and the image size as H×W; Layer 1: 3D convolution kernel size 3×3×3, stride 1, number of channels 32; Layer 2: 3D convolution + batch normalization + ReLU activation; Layer 3: temporal pooling (MaxPooling, kernel=2×1×1) to reduce temporal dimension. The morphology channel structure is the same as the color channel, the difference being the type of input feature map. After processing the two channels independently, each outputs a temporal feature tensor, denoted as F_c (color) and F_m (morphology), respectively.

[0045] S402: To enhance the model's response to keyframes in the time dimension (such as the initial flame frame or the frame of sudden smoke disturbance), this step introduces a lightweight temporal self-attention mechanism module within each channel. This module adopts the Temporal Squeeze-and-Excitation (TSE) principle and specifically includes: Temporal dimension compression (Squeeze): For each channel feature tensor Spatial average pooling is performed to obtain the time vector. ; Attention weight calculation (Excitation): Input V into a two-layer fully connected network, and output a temporal attention vector A∈R. n Normalization using the Sigmoid function is used to represent the importance weight of each frame; weighted enhancement (Recalibration): A is applied to the original feature tensor F, i.e. This mechanism enhances the features of high-importance frames. It enables the network to proactively focus on key moment frames exhibiting typical fire dynamics, thereby improving recognition sensitivity.

[0046] S403: After completing the dual-channel feature extraction, the color channel output F_c′ and the morphology channel output F_m′ are fused to generate a joint multidimensional fire feature map F_joint.

[0047] Preferably, the fusion method is feature-level weighted splicing fusion, specifically including: The feature tensors of the two channels are concatenated along the channel dimension to obtain the tensor F_concat; Feed F_concat into a Channel Attention module (such as CBAM) to automatically learn the importance weights of each channel; Output the fused tensor , where c is the number of channels after fusion.

[0048] The fused F_joint possesses the ability to express both color perturbation features and morphological perturbation features, providing rich semantic information for subsequent spatial localization.

[0049] S404: Input the fused feature map F_joint into a spatial saliency decoding network to gradually restore the spatial resolution and generate a probability map P with the same size as the original image.

[0050] The decoding network adopts a U-Net-style upsampling structure, which mainly includes: Convolutional layer group: performs semantic integration, using multiple 3×3 convolution + ReLU structures; Deconvolution layer group: progressively restores the image size (e.g., upsampling 2×); Sigmoid activation layer: Outputs a single-channel probability map , where each pixel value P(i,j)∈[0,1] represents the probability confidence that the corresponding image location is a fire area.

[0051] In a preferred embodiment of the present invention, to further determine whether there are regions with typical flame disturbance behavior in the probability map P of suspected fire areas output by the neural network, step S500 introduces a boundary stability analysis mechanism. By calculating the boundary disturbance consistency coefficient R, the structured identification of flame dynamic characteristics is achieved. Specifically, the following steps are included: S501: This step first performs region extraction processing on the probability map P of suspected fire areas in multiple consecutive frames to obtain the spatial boundaries of high-confidence suspected areas.

[0052] The specific steps are as follows: For each frame probability map Perform threshold segmentation and set a first confidence threshold. The preferred value is 0.6, which will satisfy... The pixels are considered as the pixels of the fire area, forming a binary image. ; In binary image The above contour extraction algorithm is executed, preferably using the findContours function in OpenC to extract all closed or semi-closed boundaries, thus obtaining a set of boundary contours. , where each b represents the boundary curve of a connected region; for each boundary The boundary coordinates, geometric centroid, contour length, and region area are recorded for subsequent matching analysis. After processing, a sequence of boundary contours in several consecutive frames of images can be obtained, laying the foundation for subsequent perturbation analysis.

[0053] S502: Since the flame boundary is a non-rigid target and its shape changes over time, this step uses a structural similarity algorithm to pair the boundary contours between consecutive frames to construct a boundary sequence S.

[0054] The matching process is as follows: Let the boundary set of the t-th frame be... The set of frames t+1 is ,enumerate Each boundary and Each boundary The correspondence; Define the comprehensive matching score function It consists of the following three weighted components: centroid distance error Euclidean distance between the centroids of the two boundaries; area ratio error The deviation of the ratio of region areas from 1; contour similarity. The similarity of contour shapes is measured using either the Hausdorff distance or the Fourier shape descriptor; where: The preferred weights are α1=0.4, α2=0.3, and α3=0.3. Set matching threshold When M is less than When set to 20.0, it is considered as the corresponding volume of the same boundary in consecutive frames and added to the same boundary sequence. .

[0055] After this processing, a set of multiple boundary contour sequences with temporal consistency is obtained. It is used for dynamic perturbation feature extraction.

[0056] S503: To determine whether the boundary exhibits typical flame disturbance behavior, this step structures the temporal variation behavior of the boundary into a set of disturbance feature vectors, specifically including the following three types of disturbance indicators: Deformation rate vector : Indicates the degree of variation in the area or perimeter of the outline; for each boundary Calculate its area The deformation rate is defined as follows: ; Center of gravity drift vector : Represents the trend of the boundary centroid moving over time, let The geometric center is ,but ; Edge disturbance index : Represents the edge complexity fluctuation of the contour, preferably using the boundary fractal dimension or the standard deviation of curvature change as a measure.

[0057] The above indicators are respectively composed into a time series feature vector sequence, that is , , .

[0058] S504: Finally, a comprehensive perturbation consistency coefficient R is calculated based on the above perturbation feature vector set to determine whether the boundary sequence has typical flame perturbation characteristics. The calculation method is as follows: for each perturbation vector... Calculate its standard deviation The standard deviation is normalized to avoid errors caused by different characteristic units. Define the perturbation consistency coefficient R as a weighted average: The preferred values ​​are β1=0.4, β2=0.4, and β3=0.2. Set the threshold for determining flame disturbance consistency The preferred value is 0.15, when R is less than At that time, it was believed that the boundary sequence had consistent flame perturbation and had a high probability of fire occurrence.

[0059] In a preferred embodiment of the present invention, when the boundary perturbation consistency coefficient R exceeds a set threshold... When the value is 0.15, it indicates that the region exhibits unstable and strongly disturbed boundary behavior, possessing potential flame characteristics. However, since some non-fire source disturbances (such as reflections, light spots, etc.) may also trigger R anomalies, this step further combines the high-frequency perturbation characteristics of the image's color channels, especially the frequency domain fluctuation analysis of the red channel, to generate a more reliable candidate fire source marker region F. The specific steps are as follows: S601: In this step, the color channel features C of the consecutive frame images are extracted from the image space region corresponding to the identified boundary sequence S, and the variation behavior of its red component is analyzed in detail.

[0060] The specific processing flow is as follows: for each frame of the image, obtain the red tone mapping in the HSV color space, or directly extract the red channel from the RGB image. For the spatial region covered by the boundary sequence S, in N consecutive frames, record the temporal sequence of the red channel value of each pixel (i,j) as it changes over time: Constructing a red channel tensor in the time dimension This is used for subsequent frequency domain analysis.

[0061] S602: In order to capture the rapid color change characteristics of flames (such as flickering and combustion boundary fluctuations), this step performs frequency domain analysis on the above red channel time series, extracts its high-frequency energy components, and constructs a high-frequency energy spectrum E.

[0062] The processing method is as follows: For each pixel (i, j), perform a Short-Time Fourier Transform (STFT) or Discrete Wavelet Transform (DWT) on R_seq(i, j). If STFT is used, extract the energy density of high-frequency bands (e.g., frequency ≥ 0.4 × Fs, where Fs is the frame rate) within a window length of k (preferably 5). This yields an energy spectrum E(i, j), representing the intensity of a significant high-frequency disturbance at that location in the red channel. Optionally, Gaussian smoothing is applied to spectrum E to reduce single-point errors. Finally, a two-dimensional image is formed. This is called the red high-frequency energy map, used to identify possible combustion source regions.

[0063] S603: This step analyzes the local aggregation characteristics of the energy spectrum E to screen out spatial regions with high-intensity color perturbations.

[0064] The specific implementation is as follows: Set a high-frequency energy determination threshold. Its value is dynamically set based on the statistical distribution of the training set, and is preferably the mean of the energy image plus twice the standard deviation: On the energy spectrum E, the markings satisfy... The pixels are red high-frequency abnormal points; Using connected component analysis (such as 8-neighbor connectivity), adjacent anomalous pixels are aggregated to form regions, resulting in a set of high-frequency red anomalous regions. .

[0065] S604: To ensure the consistency of spatial and temporal information, this step performs a spatial intersection operation on the high-frequency color anomaly region R_high and the boundary disturbance region S, and filters out regions that exhibit high dynamic characteristics in both color and shape as candidate fire source regions F.

[0066] Specifically, traverse each high-frequency red anomaly region. The spatial intersection of the mask region corresponding to the boundary sequence S is calculated to determine the degree of overlap. If the overlapping area exceeds a set ratio (e.g., Intersection over Union (IOU) ≥ 0.3), the region is considered a strong candidate region for fire source. All overlapping regions that meet the conditions are merged to output the final fire source marked region F. Optionally, a confidence score is assigned to each region F as one of the input features for subsequent multimodal discrimination.

[0067] This embodiment introduces a multimodal fusion discrimination mechanism in step S700, comprehensively analyzing the texture stability, perturbation directionality, and historical smoke evolution behavior of region F, thereby outputting the final fire judgment result L. Specifically, it includes the following steps: S701: This step is mainly used to determine whether the performance of region F is stable under different lighting conditions, so as to eliminate high-brightness interference caused by non-fire source phenomena such as light spots and specular reflection.

[0068] Specifically, for the image frame containing region F, three image versions are constructed: the original RGB image; a grayscale image (using a weighted grayscale transformation Y = 0.299R + 0.587G + 0.114B); and a gamma-corrected image (the gamma value γ is preferably 2.2, and the conversion formula is...). In each image version, texture features are extracted from region F, preferably using the Local Binary Pattern (LBP) method for texture encoding. The LBP histogram for region F in each image channel is calculated, and its entropy and standard deviation are determined to reflect texture stability. If the texture fluctuations in the three channels are significantly different (e.g., the cosine similarity of the LBP distribution between any two channels is less than a set threshold, such as 0.85), then the region is considered to have unstable texture, possibly a false fire source caused by light interference. This step provides a texture stability score feature value S_T for subsequent models.

[0069] S702: To distinguish between flame disturbance and boundary motion caused by mechanical disturbance or hot airflow, this step performs directional modeling of the disturbance trajectory in region F.

[0070] Specifically, in N consecutive frames of images, the outer boundary contour or centroid position of region F is traced to form a time-series trajectory. Principal component analysis (PCA) was performed on the two-dimensional trajectory data to extract the first principal direction vector. , indicating the main perturbation direction; for each frame, the perturbation vector and the main direction Statistical analysis is performed on the cosine of the angle between the two points to obtain the directional disturbance deviation sequence A_seq, and its standard deviation σ_dir is further calculated. If σ_dir is less than a set threshold (e.g., the cosine corresponding to 10 degrees is 0.984), the disturbance is judged to have strong directional consistency and is consistent with the characteristics of flame disturbance. If the deviation is large and the change is irregular, it may be due to environmental factors. This step outputs the disturbance directional consistency score S_D.

[0071] S703: This step analyzes the space surrounding area F through time backtracking to determine whether there is a smoke evolution process, and is used to help determine whether the current fire source area is accompanied by a normal fire development path.

[0072] Specifically, backtracking M frames (preferably 5 to 10 frames) from the frame containing region F, the image block sequence of its spatially adjacent regions is obtained; a smoke detection algorithm is executed on each frame image block, preferably using the following method: increasing the S channel and decreasing the V channel in HSV space; decreasing local image contrast (using local variance calculation); increasing texture blur (using high-pass filtering to decrease response); constructing a time series feature vector, and analyzing whether the above smoke indicators continuously increase; if there is a stable smoke change trend from weak to strong (e.g., the detection indicators show a monotonically increasing trend or the average rate of change is greater than a set threshold), it is considered that there is a smoke evolution trend, and a score S_S (e.g., between 0 and 1) is output.

[0073] S704: After completing the analysis of texture stability, perturbation directionality, and smoke evolution trend in region F, this step uses a multimodal fusion model for final fire identification.

[0074] Specifically, a three-dimensional feature vector is constructed: X = [S_T, S_D, S_S], representing the texture stability score, perturbation direction consistency score, and smoke evolution matching score, respectively. The feature vector is input into a pre-trained fire multimodal fusion discrimination model. The preferred model types include: multilayer perceptron (MLP), with an input layer – hidden layer – output layer and ReLU activation function; or a random forest model based on decision trees, used for scenarios with high robustness in small samples. The model output is the fire judgment result L, which can be: a binary result (L=1 indicates a fire, L=0 indicates no fire); or a continuous confidence score L∈[0,1], used by the subsequent alarm system to set dynamic thresholds.

[0075] In a preferred embodiment of the present invention, when the fire judgment result L output by the multimodal fusion discrimination model indicates that a fire has occurred, in order to achieve rapid response and subsequent source tracing of the fire, this step S800 includes the generation and transmission of an alarm signal, as well as the reconstruction and visualization annotation of the fire source trajectory and fire development path based on the candidate fire source area F, specifically including the following processing steps: S801: When the fire judgment result L is determined to be "fire state" (i.e. L=1), a structured alarm data packet containing key information is immediately generated and an alarm signal is sent through the network interface or local I / O port.

[0076] The processing flow is as follows: Construct an alarm data structure, which includes at least the following fields: alarm timestamp (T_alert), using the system clock or video frame timestamp; fire area location (F_pos), represented by the bounding box coordinates or centroid coordinates of area F; alarm level (Level), which can be automatically classified based on the area of ​​area F, confidence level L value, etc. (e.g., Level 1: small suspected area; Level 2: obvious spread trend); encapsulate the above data into a structured data packet (e.g., JSON or binary protocol format); send the alarm signal to: a remote video monitoring platform; a local intelligent terminal (e.g., a voice alarm device, fire linkage module); or a cloud-based disaster dispatch system (e.g., a city fire cloud platform) via a communication module (e.g., RS485, Ethernet, MQTT, 4G module, etc.).

[0077] S802: To assist in subsequent source tracing and ignition point location, this step starts from the current fire judgment frame. By tracing back several consecutive frames of images, the historical evolution trajectory of the candidate fire source region F can be tracked.

[0078] Specifically, set the number of backtracking frames N_back, preferably 10 to 30 frames (depending on the scene's frame rate); for each frame Candidate regions for fire sources in (k = 1 to N_back) Extract its geometric center coordinates ; Construct a time series trajectory set, This indicates the path of the fire source from before it started to the present state; Alternatively, if the area of ​​region F is extremely small or has extremely low confidence in early frames, the start time of the ignition point can be estimated. The time point when the fire first appeared was further marked.

[0079] S803: In consecutive frames following a fire detection, the system continuously tracks the spatial expansion trend of the fire source area F and constructs a dynamic fire expansion trajectory P_expand to assess the spread speed and direction. The processing method is as follows: For each frame of the image after it was determined to be a fire Extract the geometric center coordinates of region F. Area of ​​bounding box ; Contour shape (can be used to determine irregular flame boundaries); Calculate the area expansion rate Constructing the trajectory path Simultaneously, the expansion rate sequence and direction vector sequence are recorded; if the expansion rate continues to increase and is accompanied by a shift in the centroid, the direction and intensity of the fire can be dynamically determined.

[0080] S804: To assist in video scheduling and on-site handling, this step concatenates the starting path P_start and the extended trajectory P_expand to generate a complete fire development trend path P_total, and then visualizes and annotates it on the video frame.

[0081] The specific processing method is as follows: P_start and P_expand are merged into a continuous path line, and a curved structure with arrows is generated according to the trajectory direction; color coding rules are set, for example: the starting point is marked as a green dot; the fire source determination point is marked as a red flashing dot; the expansion trajectory is a red-yellow gradient line, and the darker the color, the faster the expansion; in the alarm video frame image, P_total is superimposed as a layer for drawing, and is simultaneously embedded in the alarm recording or used for real-time preview of the dispatch terminal; optionally, P_total is stored synchronously with the original image for accident playback or AI relearning training set construction.

[0082] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A fire intelligent detection method based on artificial intelligence video analysis, characterized in that: include: S100: Acquires video surveillance stream of the target area and extracts the spatial color feature vector C and morphological feature vector M of each frame image; S200: Perform temporal difference analysis on consecutive frame images to obtain the change magnitude vectors ΔC and ΔM based on C and M, and establish an initial dynamic feature set D; S300: Construct a multi-scale temporal feature map group T, wherein T is composed of a three-dimensional tensor formed by the D of N consecutive frames of images under different time windows; S400: Use a dual-channel neural network model based on attention mechanism to extract temporal features from T and output a probability map P of suspected fire areas; S500: Perform boundary stability analysis on the suspected region in P to determine whether there is a boundary sequence S with flame-shaped perturbation characteristics, and calculate its perturbation consistency coefficient R. S600: If R exceeds the set threshold, further combine the high-frequency fluctuation characteristics of the red channel in C and ΔC to generate candidate fire source marking regions F; S700: Input F into the fire multimodal fusion discrimination model, fuse the texture stability, perturbation directionality and historical smoke evolution model of region F under different lighting channels, and output the final fire judgment result L; S800: If L determines that a fire has occurred, an alarm signal is sent and the fire's starting position and development path are marked in the video frame.

2. The intelligent fire detection method based on artificial intelligence video analysis according to claim 1, characterized in that: S200 includes: S201: Perform inter-frame registration processing on multiple consecutive frames of images within a preset time window; S202: Based on the registered image frame sequence, perform frame-by-frame difference operation on the spatial color feature vector C and morphological feature vector M of each frame image to calculate the color change amplitude ΔC and morphological change amplitude ΔM between adjacent frames. The difference method includes Euclidean distance or structural similarity calculation. S203: Perform filtering and smoothing on ΔC and ΔM; S204: Combine the smoothed ΔC and ΔM to construct a dynamic feature set D, which is used to describe the evolution characteristics of suspected flame or smoke regions in a video frame sequence.

3. The intelligent fire detection method based on artificial intelligence video analysis according to claim 1, characterized in that: The S300 includes: S301: Set multiple time window scale groups n is the total number of windows, and each time window corresponds to a frame interval, which is used to extract image sequence features of different time lengths from the dynamic feature set D. S302: For each time window, extract the feature map sequence of consecutive frames from the dynamic feature set D, and stack the sequence in time order to form a primary three-dimensional tensor Ti with size Wi×H×W, where H and W are the image height and width; S303: Perform temporal normalization on each primary 3D tensor Ti to generate a temporally normalized tensor Ti′; S304: Combine the time-normalized tensors Ti′ at all scales to form a multi-scale time-series feature map group T, which is used to characterize the dynamic evolution characteristics of fire at different time granularities.

4. The intelligent fire detection method based on artificial intelligence video analysis according to claim 1, characterized in that: The S400 includes: S401: Construct a dual-channel neural network model with a parallel structure, which receives color dynamic information and morphological dynamic information from the multi-scale temporal feature map group T respectively, and independently extracts the corresponding temporal spatial feature tensors; S402: Within each channel, calculate the keyframe weights in the time dimension of the input 3D tensor. S403: Perform channel fusion processing on the feature tensors output from the two channels to generate a fused multidimensional fire feature map; S404: The fused feature map is input into the spatial saliency decoding network, and convolutional and deconvolutional layers are used to gradually restore the suspected fire area probability map P with the same size as the original image. The value of each pixel in P represents the confidence probability that the location is a fire area.

5. The intelligent fire detection method based on artificial intelligence video analysis according to claim 1, characterized in that: The S500 includes: S501: Binarize the probability map P of suspected fire areas in multiple consecutive frames, extract the boundaries of areas with confidence scores higher than the first threshold, and extract the boundary contour set B for each frame. t ; S502: On the time series, the boundary contour set B t Inter-frame matching is performed, and boundary correspondence is constructed based on the geometric centroid position of the contour, the similarity of the boundary shape, and the area change ratio to form a candidate boundary sequence S; S503: Quantify the perturbation behavior of each boundary in the boundary sequence S and construct a perturbation feature vector set, including boundary deformation rate, centroid drift vector and local edge irregularity index; S504: Calculate the perturbation consistency coefficient R based on the perturbation feature vector group. R is the normalized result of the standard deviation of the perturbation features during the boundary change over time. If R is less than the second threshold, the boundary sequence is determined to have flame perturbation consistency.

6. The intelligent fire detection method based on artificial intelligence video analysis according to claim 1, characterized in that: The S600 includes: S601: Provided that the perturbation consistency coefficient R exceeds a set threshold, extract the spatial color features C of each frame from the image region corresponding to the boundary sequence S, and focus on extracting the red channel values. Pixel distribution; S602: Perform frequency domain analysis on the time-series value sequence of the red channel in consecutive frame images, extract the high-frequency components of the red channel, and construct the red high-frequency energy spectrum E; S603: Within the region corresponding to the boundary sequence, analyze the spatial distribution of the high-frequency energy spectrum E, identify the high-frequency fluctuation clustering region, and screen out the red high-frequency anomaly region R_high; S604: Perform spatial overlap calculation on the boundary sequence S and R_high, and extract the intersection region as the candidate fire source marker region F.

7. The intelligent fire detection method based on artificial intelligence video analysis according to claim 1, characterized in that: The S700 includes: S701: Extract texture stability index for candidate fire source marking region F under multiple illumination channels, and calculate its local texture consistency score in the original image, grayscale image and gamma-corrected image respectively; S702: Based on the contour deformation trajectory of region F in consecutive frames, construct a perturbation directionality vector set, obtain the main perturbation direction through principal component analysis, and calculate the direction perturbation deviation value. S703: Call the historical smoke evolution model to perform time-reverse feature comparison on region F, assess whether there is a continuous smoke diffusion process in its surrounding areas, and output a matching score; S704: The texture stability, perturbation directionality and historical smoke matching score are taken as inputs and fed into the fire multimodal fusion discrimination model for joint decision-making, and the final fire judgment result L is output, where L is a binary result or confidence score.

8. The intelligent fire detection method based on artificial intelligence video analysis according to claim 1, characterized in that: The S800 includes: S801: When the fire judgment result L is in the established state, generate the alarm time, target area location and alarm level, and send an alarm signal; S802: Starting from the time point when the fire was determined, trace back through consecutive video frames, extract the geometric center position of the candidate fire source marking region F in each frame image, and construct the fire starting path trajectory P_start; S803: In consecutive frames after the fire occurs, dynamically track the area expansion, boundary movement trend and center drift direction of region F, and calculate the fire spread path trajectory P_expand; S804: P_start and P_expand are concatenated to generate the fire development trend path P_total, and the fire source origin, movement direction and development trend are marked with trajectory lines in the alarm video frame.

Citation Information

Patent Citations

  • Fire smoke and fire source accurate positioning method based on image recognition

    CN119784848A

  • Geological disaster early warning method based on remote sensing monitoring

    CN120452141A

  • Illegal behavior identification and early warning method, system and device for intelligent power distribution network, and medium

    CN120635525A

  • Hyperspectral-VOCs underway combined traceability method based on deep learning

    CN120635736A

  • Fire situation analysis method based on multi-dimensional data fusion

    CN120747864A

Cited By

  • Fire behavior identification method and system based on video monitoring

    CN121617049A

  • Fire identification method and system based on video monitoring

    CN121617049B

  • Intelligent edge flame and smoke identification method based on deep learning

    CN121861589A

  • Impact point positioning method and system based on machine vision

    CN121999053A

  • Intelligent fire detection method based on artificial intelligence video analysis

    CN122368924A