Intelligent fire detection and positioning method based on visual recognition

By splitting and processing video source signals and performing multi-dimensional analysis, the problem of inaccurate fire feature extraction in existing technologies has been solved, enabling accurate detection and location of fires and improving the accuracy of fire assessment and geographic location.

CN122024181BActive Publication Date: 2026-06-19BEIJING JINZHOU FIRE FIGHTING ENG CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING JINZHOU FIRE FIGHTING ENG CO LTD
Filing Date
2026-04-15
Publication Date
2026-06-19

Smart Images

  • Figure CN122024181B_ABST
    Figure CN122024181B_ABST
Patent Text Reader

Abstract

This invention relates to the field of visual recognition technology for fire detection, specifically a visual recognition-based intelligent fire detection and localization method. The method includes: acquiring continuous panoramic frame images of a scene using a video sensor array to generate a video source signal, which is then separated into a basic visual stream and an auxiliary visual stream. A steady-state feature model is obtained by performing background dynamic modeling on the basic visual stream. The auxiliary visual stream undergoes spectral transformation processing to generate an enhanced spectral response stream. Fire-related spectral features are extracted and compared point-by-point with the steady-state model to identify disturbed areas. A dual-dimensional spatial and temporal analysis is conducted on the disturbed areas. A fire score is calculated by combining spatial features such as contour, texture, and color, and temporal features such as area, shape, and positional drift. Based on the spatiotemporal information of high-scoring disturbed areas, the fire determination result and geographic location coordinates are output. This method can reduce environmental interference and improve the accuracy of fire identification and localization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of visual recognition technology for fire detection, and in particular to an intelligent fire detection and location method based on visual recognition. Background Technology

[0002] Conventional video-based fire detection technologies mostly employ a single-stream visual processing mode, directly processing panoramic continuous frame images acquired by video sensors. Fire detection is determined based on color, brightness, or texture features within a single image domain. Even technologies that incorporate spectral features do not perform separate processing of the video signal; they only perform spectral analysis and background modeling on the overall image. These technologies directly rely on single-dimensional features to mark suspected areas, lacking separate processing mechanisms for basic analysis and cross-validation. The matching of spectral features with the background model is primarily based on overall comparison.

[0003] The single-view flow processing method is easily affected by factors such as ambient lighting and clutter interference. The extraction of fire features lacks specificity, the auxiliary spectral features are not enhanced, and the overall comparison with the background steady-state model cannot accurately identify local abnormal disturbances, resulting in a large number of invalid markers. Existing technologies analyze suspected areas using only single-frame spatial features or simple temporal changes, without performing refined spatial feature quantification for disturbed areas or continuous temporal tracking of the dynamic changes in disturbed areas. This leads to potential biases in fire assessment results and insufficient geolocation accuracy.

[0004] This invention aims to solve the problem of low feature analysis accuracy caused by the lack of signal splitting in video sources. It achieves the separation of the basic visual stream and the auxiliary visual stream, and accurately marks the disturbed area by spectral enhancement of the auxiliary stream and point-by-point comparison of spectral features with the steady-state model. At the same time, it solves the problem of the single dimension of disturbance area analysis by performing spatial domain multi-feature calculation and temporal domain dynamic trajectory tracking on the disturbance area, and completes fire situation determination and location based on the multi-dimensional analysis results. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of existing technologies by proposing a visual recognition-based intelligent fire detection and location method.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: an intelligent fire detection and location method based on visual recognition, comprising:

[0007] Video source signals are generated from continuous frame images containing panoramic views of the scene acquired by a video sensor array.

[0008] The basic visual stream for preliminary analysis and the auxiliary visual stream for cross-validation are separated from the video source signal.

[0009] The basic visual flow is scanned frame by frame, and a steady-state feature model of the scene is generated through dynamic background modeling.

[0010] Perform spectral transformation on the auxiliary visual stream to generate an enhanced spectral response stream;

[0011] Features matching the preset potential fire spectral range are extracted from the enhanced spectral response stream to generate a potential fire spectral feature set. The potential fire spectral feature set is then compared point by point with the steady-state feature model to mark the perturbation regions that exceed the steady-state features and generate a primary perturbation region list.

[0012] Based on the list of primary disturbance regions, a multi-stage analysis of the video source signal is initiated, which includes a spatial domain analysis stage and a temporal domain analysis stage.

[0013] In the spatial domain analysis phase, the boundary contour, internal texture complexity, and color statistical distribution of each primary perturbation region are calculated.

[0014] In the temporal analysis phase, the area change, shape evolution path, and positional drift trajectory of each primary perturbation region are tracked across multiple frames of images.

[0015] Based on the combined spatial domain analysis results and temporal domain analysis results, the fire score for each primary disturbance area is calculated.

[0016] Based on the spatiotemporal information of the primary disturbance area where the fire score exceeds the decision threshold, a fire determination result and geographic location coordinates are generated, and control commands are output.

[0017] As a further aspect of the present invention, the step of scanning the basic visual stream frame by frame and generating a steady-state feature model of the scene through background dynamic modeling includes:

[0018] Starting from the beginning of the basic visual flow, select a series of consecutive frames as reference frames for background modeling;

[0019] Extract the color value sequence of each pixel in the background modeling reference frame across consecutive frames;

[0020] For each pixel, a Gaussian mixture model is fitted to its color value sequence to generate a set of multimodal steady-state Gaussian components for each pixel.

[0021] The set of multimodal steady-state Gaussian components of a pixel is weighted and evaluated, and the steady-state Gaussian component with the highest weight is defined as the steady-state feature Gaussian component of the pixel.

[0022] Collect the Gaussian components of the steady-state features of all pixels to form a two-dimensional steady-state feature matrix;

[0023] Calculate the similarity of the Gaussian components of the steady-state features of adjacent pixels in the color space in the two-dimensional steady-state feature matrix, and construct the boundary of the steady-state feature region based on the similarity;

[0024] Based on the boundaries of the steady-state feature regions, adjacent pixels with similar steady-state Gaussian components are grouped into steady-state feature blocks;

[0025] Each steady-state feature block is assigned a unique feature block identifier, and the average color value and spatial covariance matrix corresponding to the feature block identifier are recorded, which together form the steady-state feature model.

[0026] As a further aspect of the present invention, the step of performing spectral transformation processing on the auxiliary visual stream to generate an enhanced spectral response stream includes:

[0027] Infrared spectral images and short-wave infrared spectral images are separated from the raw auxiliary visual stream acquired by the multispectral camera.

[0028] Histogram equalization is performed on the infrared spectral band image to enhance the overall image contrast and generate the first processed spectral image.

[0029] Adaptive gamma correction is applied to shortwave infrared spectral images to highlight the energy characteristics of preset bands and generate second-processed spectral images.

[0030] The first processed spectral image and the second processed spectral image are then fused at the pixel level to generate a fused spectral image;

[0031] A filter based on a preset fire spectral reference is applied to the fused spectral image, the preset fire spectral reference being generated based on a standard flame spectral database;

[0032] Calculate the response intensity of each pixel in the fused spectral image with the preset fire spectral reference to obtain a preliminary spectral response map;

[0033] Multi-scale spatial morphological operations are performed on the preliminary spectral response map to remove isolated noise response points and connect adjacent strong response regions to form connected regions.

[0034] Extract the geometric center coordinates, bounding rectangle, and response intensity peak of each connected region to form a list of spectral feature descriptors;

[0035] Based on the list of spectral feature descriptors, the corresponding response regions are marked on each frame of the original auxiliary visual stream to generate the enhanced spectral response stream.

[0036] As a further aspect of the present invention, the step of generating a potential fire spectral feature set and comparing the potential fire spectral feature set with the steady-state feature model point by point, and marking the perturbation region exceeding the steady-state features, includes:

[0037] Iterate through each response region defined by the list of spectral feature descriptors in the enhanced spectral response stream;

[0038] Locate the pixel position in the basic visual flow that corresponds to the geometric center coordinates of the response region;

[0039] From the steady-state feature model, read the mean and variance of the steady-state feature Gaussian components corresponding to the pixel position;

[0040] Read the actual color value of the pixel location from the current frame of the basic visual stream;

[0041] Calculate the probability that the actual color value falls within the color distribution range defined by the corresponding steady-state Gaussian component, and generate a steady-state conformity value;

[0042] If the steady-state compliance value is lower than the set steady-state deviation threshold, then the pixel position is determined to be a steady-state deviation point;

[0043] Collect all steady-state deviations and mark the response regions to which they belong as suspicious regions;

[0044] For all suspicious regions in the basic visual flow, the steady-state conformity value is repeatedly calculated, and the number of steady-state deviation points and the proportion of the total area in each suspicious region are counted.

[0045] The overall perturbation intensity of the suspected region is obtained by weighted calculation by combining the peak response intensity of the suspected region in the enhanced spectral response stream.

[0046] Suspicious regions whose overall disturbance intensity exceeds the disturbance threshold are filtered to form the primary disturbance region list. The entries in the primary disturbance region list include the boundary information and overall disturbance intensity of each region.

[0047] As a further aspect of the present invention, the calculation of the boundary contour, internal texture complexity, and color statistical distribution of each primary perturbation region includes:

[0048] For each primary perturbation region, extract the set of pixels from its boundary information;

[0049] Perform convex hull calculation on the set of pixels to generate a convex polygonal contour of the primary perturbation region, and record the coordinates of each vertex of the convex polygonal contour;

[0050] Based on the convex polygon profile, the ratio of its area to its perimeter is calculated as a profile compactness feature.

[0051] Within the boundaries of the primary perturbation region, extract grayscale image sub-blocks;

[0052] The local binary pattern operator is applied to the image sub-blocks, and the histograms of different local binary pattern codes are statistically analyzed. The entropy value of the histogram is calculated as an internal texture complexity feature.

[0053] Within the boundary of the primary perturbation region, the mean, variance, skewness, and kurtosis of all pixels in multiple color channels are statistically analyzed to form a color statistical distribution vector.

[0054] The contour compactness feature, internal texture complexity feature, and color statistical distribution vector are combined to generate the spatial feature vector of the primary perturbation region.

[0055] As a further aspect of the present invention, the tracking of the area change, shape evolution path, and positional drift trajectory of each primary perturbation region in multiple frames of images includes:

[0056] In the subsequent frames of the base visual stream containing the primary perturbation region, each primary perturbation region is dynamically tracked.

[0057] In the current frame, starting from the center position of the primary perturbation region determined in the previous frame, template matching is performed within a preset neighborhood range to determine the approximate position of the primary perturbation region in the current frame.

[0058] At the approximate location, the color statistical distribution vector of the previous frame of the primary perturbation region is used as a reference to perform region growing segmentation to obtain the segmented region of the current frame.

[0059] Calculate the area of ​​the segmented region in the current frame and compare it with the area in the previous frame to calculate the rate of change of area.

[0060] Extract the convex polygon contour of the segmented region in the current frame, calculate its Jaccard similarity coefficient with the convex polygon contours of the previous several frames, and describe the shape evolution path.

[0061] Record the geometric center coordinates of the segmented region in the current frame, calculate the Euclidean distance between it and the geometric center coordinates of the previous frame, and describe the inter-frame position drift.

[0062] Smoothing filters are applied to the area change rate sequence, Jaccard similarity coefficient sequence, and position drift distance sequence of consecutive frames to remove noise points;

[0063] Calculate the mean and variance of the smoothed area change rate sequence, calculate the overall descent slope of the Jaccard similarity coefficient sequence, and calculate the cumulative drift of the position drift distance sequence.

[0064] The mean and variance of the area change rate sequence, the overall descent slope of the Jaccard similarity coefficient sequence, and the cumulative drift of the position drift distance sequence are combined to generate the temporal evolution feature vector of the primary perturbation region.

[0065] As a further aspect of the present invention, the calculation of the fire score for each primary disturbance area by integrating the spatial domain analysis results and the temporal domain analysis results includes:

[0066] For each primary perturbation region, its spatial feature vector and temporal evolution feature vector are concatenated to form a combined feature vector;

[0067] Retrieve from the historical fire case database several historical cases that are closest in feature space distance to the current combined feature vector;

[0068] Read the final fire situation assessment result tags corresponding to historical cases;

[0069] Based on the distance between historical cases and the current combined feature vector, each historical case is assigned a different weight, with the weight increasing as the distance increases.

[0070] The final fire situation determination labels of historical cases are weighted and voted on. The ratio of the total weight of the vote result "it is a fire" to the total weight of all votes is calculated to obtain the initial fire situation probability based on case matching.

[0071] The average value of the comprehensive disturbance intensity, internal texture complexity features, and area change rate sequence of the current primary disturbance region is nonlinearly combined and mapped into a correction factor.

[0072] The correction factor is used to adjust the initial fire probability based on case matching to obtain the corrected fire score.

[0073] As a further aspect of the present invention, the generation of fire determination results and geographic location coordinates includes:

[0074] All primary disturbance areas whose fire scores exceed the preset fire judgment threshold are selected;

[0075] For each primary perturbation region that meets the threshold condition, extract the set of vertex coordinates of the convex polygon contour from its spatial feature vector.

[0076] The vertex coordinate set is mapped from the image coordinate system to the actual geographic coordinate system through a pre-calibrated homography transformation matrix to obtain the polygon boundary coordinates of the fire area in geographic space.

[0077] For the same suspected fire source, if there are primary disturbance regions that meet the threshold conditions in multiple consecutive frames, they are considered as different time observations of the same fire event.

[0078] The geographical polygon boundary coordinates calculated in different frames of the fire event are fused together, and the union of the polygon coordinates is taken as the final geographical location of the fire event.

[0079] Generate a fire event identifier for each individual final location geographic area;

[0080] The fire event identifier, the boundary coordinates of the final geographical area, the timestamp of the first detection, and the corresponding average fire score are bound together to form a complete fire determination and location record.

[0081] The fire detection and location records are encapsulated into a structured data format and output as the core content of the control commands.

[0082] As a further aspect of the present invention, after outputting the control command, a feedback learning phase is performed, the feedback learning phase including:

[0083] After the control command triggers an external response, it continuously collects subsequent monitoring video segments of a predetermined length;

[0084] In subsequent monitoring video segments, the actual development of the fire is confirmed manually or through more reliable independent sensors, including whether the fire is real and the actual extent of its spread, and a fire truth value label is generated.

[0085] From the beginning of subsequent monitoring video segments, extract the actual evolution data of the regions in the list of primary disturbance regions within this time period, including actual spatial feature changes and temporal evolution trajectories;

[0086] The actual evolution data is compared with the initially generated spatial feature vector and temporal evolution feature vector to calculate the feature prediction bias;

[0087] The feature prediction bias, the initial fire score, and the fire truth label are correlated to form a feedback learning sample.

[0088] Feedback learning samples are stored in a historical fire case database and used to update the weight allocation rules or nonlinear combination function parameters based on case matching in the fire scoring calculation model.

[0089] As a further aspect of the present invention, a scene adaptive modeling phase is also included, which is executed during system initialization. The scene adaptive modeling phase includes:

[0090] In a fire-free monitoring scenario, video data from a complete natural day is collected as a sample for scene modeling.

[0091] Analyze the scene modeling samples to identify periodically occurring interference sources in the scene, including vehicle lights, the on of lighting fixtures, and moving sunlight spots;

[0092] Extract the visual feature patterns of each periodic interference source and their temporal occurrence patterns to generate an interference source feature-time pattern library;

[0093] In subsequent real-time detection, when the spatial characteristics and occurrence time of the primary disturbance region identified in the basic visual flow and the auxiliary visual flow match a certain pattern in the interference source feature-time pattern library with a preset interference matching threshold, the comprehensive disturbance intensity of the primary disturbance region will be significantly reduced, or it will be directly removed from the list of primary disturbance regions.

[0094] In the scene modeling samples, the natural fluctuation range of steady-state characteristic model parameters under different weather conditions was statistically analyzed;

[0095] The natural fluctuation range is used as the dynamic threshold for updating the steady-state feature model. In real-time detection, the steady-state feature model is only slowly updated when the color change of a pixel exceeds this dynamic threshold, thus avoiding the model being contaminated by brief interference.

[0096] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0097] The video source signal is separated into a basic visual stream for preliminary analysis and an auxiliary visual stream for cross-validation. The basic visual stream is scanned frame-by-frame, and a steady-state feature model of the scene is generated through dynamic background modeling. The auxiliary visual stream undergoes spectral transformation to generate an enhanced spectral response stream. Features matching the preset potential fire spectral range are extracted from the enhanced spectral response stream to generate a potential fire spectral feature set. This potential fire spectral feature set is then compared point-by-point with the steady-state feature model, marking perturbation regions exceeding the steady-state features and generating a primary perturbation region list. This dual-visual-stream splitting process distinguishes between the basic analysis and cross-validation channels. Spectral transformation enhances fire-related spectral response features, and point-by-point comparison eliminates conventional steady-state areas of the scene, accurately locating local areas with abnormal features, avoiding feature confusion issues caused by overall image analysis, and reducing anomaly marking caused by non-fire factors.

[0098] A multi-stage analysis process, combining spatial and temporal domain analysis, is initiated based on a list of primary disturbance areas. In the spatial domain analysis phase, the boundary contour, internal texture complexity, and color statistical distribution of each primary disturbance area are calculated. In the temporal domain analysis phase, the area change, shape evolution path, and positional drift trajectory of each primary disturbance area across multiple image frames are tracked. The fire severity score for each primary disturbance area is calculated by combining the spatial and temporal analysis results. Based on the spatiotemporal information of areas where the fire severity score exceeds the decision threshold, a fire judgment result and geographic location coordinates are generated, and control commands are output. Spatial domain multi-feature calculation can fully present the static fire severity characteristics of the disturbance areas, while continuous temporal tracking can reconstruct the dynamic changes of the disturbance areas. Combining static features with dynamic trajectories can refine the fire severity discrimination criteria, distinguishing between fire disturbances and conventional environmental interference. Based on spatiotemporal information, the geographic coordinates of the fire area can be directly determined, and corresponding control commands are output. Attached Figure Description

[0099] Figure 1 This is a flowchart of an intelligent fire detection and location method based on visual recognition as described in this invention;

[0100] Figure 2 A flowchart for enhancing the spectral response stream generation method;

[0101] Figure 3 The curve for fire situation scoring evolution and classification determination;

[0102] Figure 4 Box plot of feature space distance versus weight distribution;

[0103] Figure 5 The performance curve of the feedback learning-gradient descent model is optimized. Detailed Implementation

[0104] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0105] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0106] See Figure 1 This invention provides an intelligent fire detection and location method based on visual recognition, the overall implementation scheme of which is as follows:

[0107] A series of frame images encompassing the panoramic view of the monitored scene are acquired using a video sensor array to generate the raw video source signal. From this video source signal, a basic visual stream for core analysis and an auxiliary visual stream for cross-validation are separated. The basic visual stream is scanned frame-by-frame, and a steady-state feature model reflecting the scene's normal state is constructed through background dynamic modeling. Simultaneously, the auxiliary visual stream undergoes spectral transformation processing to generate an enhanced spectral response stream highlighting specific spectral features. Features matching the preset potential fire spectral range are extracted from the enhanced spectral response stream to form a potential fire spectral feature set. This feature set is then compared point-by-point with the steady-state feature model to identify regions deviating from the steady-state features, generating a list of primary disturbance regions. Based on this list, a multi-stage analysis of the video source signal is initiated, including spatial and temporal domain analysis stages. In the spatial domain analysis stage, the boundary contour geometry, internal texture complexity, and multi-channel color statistical distribution of each primary disturbance region are calculated. In the temporal domain analysis stage, the area change dynamics, shape evolution path, and positional drift trajectory of each primary disturbance region are tracked across multiple consecutive frames of images. Based on the combined spatial and temporal domain analysis results, a quantitative fire score is calculated for each primary disturbance area. Areas with fire scores exceeding a preset decision threshold are selected, and fire determination results and corresponding geographic coordinates are generated based on their spatiotemporal information, along with corresponding control commands.

[0108] In one embodiment of the present invention, the process of achieving dynamic background modeling to obtain a steady-state feature model in the processing of the basic visual flow is as follows: A continuous image sequence is selected from the beginning of the basic visual flow as a reference frame for background modeling. The color value sequence of each pixel in the reference frame across multiple consecutive frames is extracted. For each independent pixel, a Gaussian mixture model is applied to its color value sequence to generate a set of multimodal steady-state Gaussian components representing the various stable states that the pixel may be in. The weights of each Gaussian component in this set are evaluated, and the component with the highest weight is defined as the steady-state feature Gaussian component of that pixel. The steady-state feature Gaussian components of all pixels are aggregated to form a two-dimensional steady-state feature matrix corresponding to the image size. The similarity measure of the steady-state feature Gaussian components of adjacent pixels in the matrix in the color space is calculated, and the boundary of the steady-state feature region is constructed based on this similarity. According to these boundaries, spatially adjacent pixels with similar steady-state feature Gaussian components are aggregated to form a larger steady-state feature block. Each steady-state feature block is assigned a unique feature block identifier, and the average color value and spatial covariance matrix corresponding to the identifier are recorded. These pieces of information together constitute a complete steady-state feature model.

[0109] In the processing of assisted visual flow, the process of generating enhanced spectral response flow is as follows: (See...) Figure 2From the raw auxiliary visual stream data acquired by a multispectral camera, infrared and short-wave infrared spectral images are separated. Histogram equalization is performed on the infrared spectral image to enhance overall contrast, resulting in a first processed spectral image. Adaptive gamma correction is applied to the short-wave infrared spectral image to highlight the spectral energy characteristics of a specific preset band, generating a second processed spectral image. The first and second processed spectral images are then pixel-level weighted fusion to produce a fused spectral image. A filter based on a preset flame spectral reference, derived from a standard flame spectral database, is applied to this fused spectral image. The response intensity of each pixel in the fused spectral image to the preset fire spectral reference is calculated, resulting in a preliminary spectral response map. Multi-scale spatial morphological operations are performed on this preliminary spectral response map to remove isolated noise response points and connect adjacent strong response regions, forming several connected regions. The geometric center coordinates, bounding rectangle, and response intensity peak of each connected region are extracted to form a list of spectral feature descriptors. Based on this list of spectral feature descriptors, the corresponding response regions are marked on each frame of the original auxiliary visual stream, ultimately generating an enhanced spectral response stream.

[0110] In practical implementation, a video sensor group deployed in a forest fire monitoring scenario is used. This group consists of a high-definition visible light camera and a multispectral camera. The high-definition visible light camera outputs an RGB color video stream as the basic visual stream, while the multispectral camera outputs an image sequence containing near-infrared and short-wave infrared bands as an auxiliary visual stream. In the basic visual stream, the number of background modeling reference frames is 200 consecutive images. The color value sequence of each pixel in the background modeling reference frames is extracted from the consecutive frames. This color value sequence is a set of pixel values ​​in the RGB color space arranged chronologically. For each pixel, a Gaussian mixture model is fitted to its color value sequence using the expectation-maximization algorithm. This generates a set of multimodal steady-state Gaussian components for each pixel. Each Gaussian component in the multimodal steady-state Gaussian component set is parameterized by weights, a mean vector, and a covariance matrix. The weights of the multimodal steady-state Gaussian component set of each pixel are evaluated, and the steady-state Gaussian component with the highest weight is defined as the steady-state feature Gaussian component of the pixel. The steady-state Gaussian components of all pixels are collected to form a two-dimensional steady-state feature matrix. Each element of the two-dimensional steady-state feature matrix stores the parameters of the steady-state Gaussian component of the corresponding pixel. The similarity of the steady-state Gaussian components of adjacent pixels in the two-dimensional steady-state feature matrix within the color space is calculated. The similarity is based on Mahalanobis distance, which measures the difference between two Gaussian distributions. Boundaries of steady-state feature regions are constructed based on the similarity. When the Mahalanobis distance between the steady-state Gaussian components of adjacent pixels is less than a similarity threshold, the pixels are considered to belong to the same steady-state feature region. Based on the boundaries of the steady-state feature regions, adjacent pixels with similar steady-state Gaussian components are grouped into steady-state feature blocks. A unique feature block identifier is assigned to each steady-state feature block, and the average color value and spatial covariance matrix corresponding to the feature block identifier are recorded, together forming the steady-state feature model.

[0111] In practical implementation, the generation of the auxiliary visual stream involves multispectral cameras deployed in the same forest fire monitoring scenario, simultaneously acquiring image sequences containing infrared and shortwave infrared bands. From the raw auxiliary visual stream acquired by the multispectral cameras, infrared and shortwave infrared spectral images are separated. Histogram equalization is performed on the infrared spectral images to expand the image's grayscale distribution, generating a first processed spectral image. Adaptive gamma correction is applied to the shortwave infrared spectral images, adjusting the gamma value based on local image brightness to highlight the energy characteristics of a preset band, generating a second processed spectral image. The first and second processed spectral images are then pixel-level weighted fusion, with the first processed spectral image having a weight coefficient of 0.6 and the second processed spectral image having a weight coefficient of 0.4. A filter based on a preset fire spectral reference is applied to the fused spectral image. This preset fire spectral reference is generated based on a standard flame spectral database, which contains characteristic spectral curves of different substances burning in the infrared and shortwave infrared bands. The response intensity of each pixel in the fused spectral image is calculated relative to a preset fire spectral reference, resulting in a preliminary spectral response map. The response intensity is obtained by calculating the cosine similarity between the pixel's spectral vector and the preset fire spectral reference vector. Multi-scale spatial morphological operations are performed on the preliminary spectral response map, including opening operations (erosion followed by dilation) and closing operations (dilation followed by erosion). Isolated noise response points are removed, and neighboring strong response regions are connected to form connected regions. The geometric center coordinates, bounding rectangle, and peak response intensity of each connected region are extracted to form a list of spectral feature descriptors. Based on this list, the corresponding response regions are marked on each frame of the original auxiliary visual stream, generating an enhanced spectral response stream.

[0112] In some embodiments, the weight evaluation of the Gaussian components of the steady-state features follows specific rules. The weight evaluation compares the weight parameters of each component in the Gaussian mixture model, where each weight parameter reflects the probability that a data point in the color value sequence belongs to that component. The steady-state Gaussian component with the highest weight represents the most frequently presented visual state of the pixel. The construction of the two-dimensional steady-state feature matrix completes the mathematical description of the normal state of the monitoring scene. The boundary delineation of the steady-state feature region is based on the continuity of the steady-state features of the pixels. The formula for calculating the Mahalanobis distance is:

[0113]

[0114] in: Represents the Gaussian components of the steady-state features of pixels i and j. and Mahalanobis distance between them and These are Gaussian components. and The mean vector, This involves merging the covariance matrices. Merging steady-state feature blocks reduces model complexity, and feature block identification facilitates rapid subsequent retrieval. It's understandable that multispectral data processing enhances sensitivity to flame spectral features. Infrared spectral bands are sensitive to thermal radiation, while short-wave infrared spectral bands are sensitive to combustion products of specific chemical substances. Histogram equalization enhances the global contrast of the first-processed spectral image, making temperature differences more apparent. Adaptive gamma correction optimizes the local contrast of the second-processed spectral image, highlighting the spectral features of combustion products. Weighted fusion comprehensively utilizes information from different bands, and a filter with a pre-set fire spectral benchmark provides a standard comparison template. Multi-scale spatial morphological operations eliminate small-particle noise and connect real fire areas, while a list of spectral feature descriptors quantifies and records the attributes of each suspected area.

[0115] In one embodiment of the present invention, the process of generating a potential fire spectral feature set and comparing it with a steady-state feature model to mark perturbation regions is as follows: Traverse each response region defined by the spectral feature descriptor list in the enhanced spectral response stream. In the corresponding frame of the basic visual stream, find the pixel position that matches the geometric center coordinates of the response region. From the constructed steady-state feature model, read the mean and variance parameters of the steady-state feature Gaussian component corresponding to the pixel position. Simultaneously, read the actual color value of the pixel position from the current frame of the basic visual stream. Calculate the probability that the actual color value falls within the color distribution interval defined by its corresponding steady-state feature Gaussian component; this probability value is the steady-state conformity value. If the calculated steady-state conformity value is lower than a pre-set steady-state deviation threshold, the pixel position is determined to be a steady-state deviation point. Collect all steady-state deviation points and mark their respective response regions as suspicious regions. Repeat the steady-state conformity value calculation for pixels in all suspicious regions in the basic visual stream, and count the number of steady-state deviation points in each suspicious region and their proportion of the total area of ​​the region. By combining the existing peak response intensity of the suspected region in the enhanced spectral response stream, a comprehensive perturbation intensity value is obtained through weighted calculation. Suspicious regions whose comprehensive perturbation intensity values ​​exceed a preset perturbation threshold are selected to form a primary perturbation region list. Each entry in this list contains the boundary information of the corresponding region and the calculated comprehensive perturbation intensity.

[0116] The process of performing spatial domain analysis on the primary perturbation region to calculate its boundary contour, internal texture complexity, and color statistical distribution is as follows: For each region in the list of primary perturbation regions, extract the set of all pixels constituting that region from its boundary information. Perform convex hull calculation on this set of pixels to generate a convex polygon contour of the region, and record the coordinates of all vertices of this contour. Based on the obtained convex polygon contour, calculate its area-to-perimeter ratio as a feature describing the contour compactness. Within the boundary range of the primary perturbation region, extract grayscale image sub-blocks. Apply the local binary pattern operator to this image sub-block, statistically analyze the generated local binary pattern encoded histogram, and calculate the entropy value of the histogram; this entropy value is the internal texture complexity feature. Within the boundary of the same region, statistically analyze the mean, variance, skewness, and kurtosis of all pixels in multiple color channels to form a color statistical distribution vector. Combine the contour compactness feature, internal texture complexity feature, and color statistical distribution vector to generate the spatial feature vector of the primary perturbation region.

[0117] In practice, each response region defined by the spectral feature descriptor list in the enhanced spectral response stream is traversed. This list contains the geometric center coordinates, bounding rectangles, and peak response intensities of multiple response regions. The pixel location corresponding to the geometric center coordinates of the response region is located in the basic visual stream, which is an RGB color video stream captured by a high-definition visible light camera. From the steady-state feature model, the mean and variance of the steady-state Gaussian component corresponding to the pixel location are read. The steady-state Gaussian component is the component with the highest weight in the Gaussian mixture model. From the current frame of the basic visual stream, the actual color value of the pixel location is read; this value is the RGB three-channel value. The probability that the actual color value falls within the color distribution range defined by the corresponding steady-state Gaussian component is calculated, generating a steady-state conformance value, which is a probability value between 0 and 1. If the steady-state conformance value is lower than a set steady-state deviation threshold (set to 0.3), the pixel location is determined to be a steady-state deviation point. All steady-state deviation points are collected, and their respective response regions are marked as suspicious regions. Steady-state conformity values ​​are repeatedly calculated for pixels within all suspicious regions in the basic visual stream, and the number of steady-state deviation points and their proportion of the total area within each suspicious region are statistically analyzed. The overall perturbation intensity of the suspicious region is calculated by combining the peak response intensity of the suspicious region with that of the enhanced spectral response stream. Suspicious regions with overall perturbation intensities exceeding a perturbation threshold (set to 0.65) are selected to form a primary perturbation region list. Each entry in the primary perturbation region list includes the boundary information and overall perturbation intensity of each region.

[0118] In practice, the process of calculating the boundary contour, internal texture complexity, and color statistical distribution of each primary perturbation region is as follows: For each primary perturbation region, a set of pixels is extracted from the boundary information of the region, which defines the pixel position of the region in the image. Convex hull calculation is performed on the pixel set using the Graham scan algorithm to generate a convex polygon contour of the primary perturbation region, and the coordinates of each vertex of the convex polygon contour are recorded. Based on the convex polygon contour, the ratio of the area to the perimeter of the convex polygon contour is calculated as the contour compactness feature. Within the boundary of the primary perturbation region, grayscale image sub-blocks are extracted. Grayscale conversion uses a weighted average method to convert RGB colors into single-channel grayscale values. A local binary mode operator is applied to the image sub-blocks, using a configuration with a radius of 1 pixel and 8 neighboring points. Histograms of different local binary mode codes are statistically analyzed, and the entropy value of the histogram is calculated as the internal texture complexity feature. Within the boundary of the primary perturbation region, the mean, variance, skewness, and kurtosis of all pixels in multiple color channels (red, green, and blue) are statistically analyzed to form a color statistical distribution vector. The contour compactness feature, internal texture complexity feature, and color statistical distribution vector are then combined to generate the spatial feature vector of the primary perturbation region.

[0119] In some embodiments, the overall disturbance intensity is calculated using a weighted summation formula. The formula for calculating the overall disturbance intensity is:

[0120]

[0121] in: Indicates the overall disturbance intensity. This indicates the number of steady-state deviations within the suspicious region. This represents the total number of pixels within the suspicious area. This indicates the peak response intensity of the suspicious region in the enhanced spectral response stream. and It's the weighting coefficient, the weighting coefficient. Set to 0.7, weighting coefficient Set to 0.3. The steady-state conformance value is calculated based on the probability density function of a multidimensional Gaussian distribution. Contour compactness reflects the regularity of a region's shape; circular regions have higher contour compactness values, while irregularly shaped regions have lower values. The histogram entropy value of the local binary mode operator quantifies the texture clutter of image sub-blocks; flame regions typically have higher internal texture complexity values. The color statistical distribution vector captures the color moment information of a region in the red, green, and blue channels.

[0122] In one embodiment of the present invention, during the temporal analysis stage, tracking the change process of the primary perturbation region in multiple frames of images involves: dynamically tracking each region in several consecutive frames of the basic visual flow following the current frame containing the primary perturbation region. In the processing of the current frame, starting from the center position of the primary perturbation region determined in the previous frame, template matching is performed within a preset neighborhood search range to determine the approximate position of the region in the current frame. At this approximate position, using the color statistical distribution vector of the region in the previous frame as a color reference, a region growing image segmentation algorithm is executed to obtain the segmentation result of the region in the current frame. The pixel area of ​​the segmented region in the current frame is calculated and compared with the area of ​​the region in the previous frame to calculate the area change rate. The convex polygon contour of the segmented region in the current frame is extracted, and its Jaccard similarity coefficient with the contours of several previous frames is calculated to describe the shape evolution path. The geometric center coordinates of the segmented region in the current frame are recorded, and the Euclidean distance between it and the center coordinates of the previous frame is calculated to describe the inter-frame positional drift. The area change rate sequence, Jaccard similarity coefficient sequence, and position drift distance sequence obtained from multiple consecutive frames are smoothed and filtered to remove anomalous noise points. The mean and variance of the smoothed area change rate sequence are calculated, the overall descent slope of the Jaccard similarity coefficient sequence is calculated, and the cumulative drift amount of the position drift distance sequence is calculated. The mean and variance of the area change rate, the overall descent slope of the Jaccard similarity coefficient, and the cumulative drift amount are combined to generate the temporal evolution feature vector of the primary perturbation region.

[0123] In the specific implementation, in the subsequent consecutive frames of the basic visual flow containing the primary perturbation region, dynamic tracking is performed on each primary perturbation region, with the number of consecutive frames set to 20. In the current frame, starting from the center position of the primary perturbation region determined in the previous frame, template matching is performed within a preset neighborhood range. The preset neighborhood range is a square region with a side length of 41 pixels centered at the starting point. The approximate position of the primary perturbation region in the current frame is determined. At the approximate position, using the color statistical distribution vector of the primary perturbation region in the previous frame as a reference, region growing segmentation is performed. The similarity threshold for region growing segmentation is set to 15 color space units to obtain the segmented region of the current frame. The area of ​​the segmented region in the current frame is calculated and compared with the area of ​​the previous frame, and the area change rate is calculated. The convex polygon contour of the segmented region in the current frame is extracted, and the Jaccard similarity coefficient between the convex polygon contour and the convex polygon contour of the previous few frames is calculated. The previous few frames refer to the 5 consecutive frames before the current frame, describing the shape evolution path. The geometric center coordinates of the segmented region in the current frame are recorded, and the Euclidean distance between the geometric center coordinates of the current frame and the geometric center coordinates of the previous frame is calculated to describe the inter-frame positional drift. The area change rate sequence, Jaccard similarity coefficient sequence, and position drift distance sequence of consecutive frames are smoothed using a median filter with a window size of 3 to remove noise points. The mean and variance of the smoothed area change rate sequence are calculated, the overall descent slope of the Jaccard similarity coefficient sequence is calculated, and the cumulative drift of the position drift distance sequence is calculated. The mean and variance of the area change rate sequence, the overall descent slope of the Jaccard similarity coefficient sequence, and the cumulative drift of the position drift distance sequence are combined to generate the temporal evolution feature vector of the primary perturbation region.

[0124] In practical implementation, dynamic tracking, template matching, and region growth segmentation ensure continuous capture of the same potential fire area. The area change rate reflects the temporal dynamics of the area size; flame spread typically results in a positive and large area change rate. The shape evolution path is characterized by a Jaccard similarity coefficient sequence; flame shapes are irregular and change rapidly, and the Jaccard similarity coefficient sequence usually shows a decreasing trend. Inter-frame positional drift describes the movement of the area center; due to the influence of smoke and airflow, the flame area may experience slight positional drift. Median filtering eliminates instantaneous outliers in the measurement. The average value of the area change rate sequence represents the average growth trend of the area, while the variance of the area change rate sequence represents the degree of fluctuation in area change. The overall decreasing slope of the Jaccard similarity coefficient sequence quantifies the decay rate of shape stability. The cumulative drift of the positional drift distance sequence represents the total displacement of the area center during the tracking period. The temporal evolution feature vector integrates dynamic information from three dimensions: area, shape, and position. It can be understood that template matching is performed within a preset neighborhood, which is set based on the maximum possible motion speed of objects between adjacent frames. Region growing segmentation uses the color statistical distribution vector of the previous frame as a seed feature to ensure continuity in color features across the segmented regions. The area change rate is calculated by subtracting the area of ​​the segmented region in the previous frame from the area of ​​the segmented region in the current frame, and then dividing by the area of ​​the segmented region in the previous frame. The Jaccard similarity coefficient is calculated based on the intersection-union ratio of the convex polygon contour. The positional drift distance is the straight-line distance between two points. Smoothing filtering improves the quality of the sequence data. The generation of temporal evolution feature vectors provides a quantitative description of the dynamic behavior of the primary perturbation regions.

[0125] In some embodiments, the formula for calculating the Jaccard similarity coefficient is:

[0126]

[0127] in: The convex polygon outline representing the segmented region of the current frame t. The convex polygon outline of the segmented region in the previous k-th frame The Jaccard similarity coefficient between them, symbol This represents the pixel area of ​​the region enclosed by the outline. This represents the pixel area of ​​the contour intersection region. The pixel area represents the region of the contour union. The overall descent slope of the Jaccard similarity coefficient sequence is obtained by linearly fitting the sequence values. The cumulative drift is the arithmetic sum of the inter-frame positional drift distances in each of several consecutive frames. Template matching can use a normalized cross-correlation method. The similarity threshold for region growing segmentation is calculated separately for the red, green, and blue channels. Smoothing filtering can also use mean filtering. In a specific implementation, for an example of the area change rate, Jaccard similarity coefficient, and positional drift distance of a tracked primary perturbation region over 5 consecutive frames, see Table 1.

[0128] Table 1: Example table of area change rate, Jaccard similarity coefficient, and location drift distance

[0129] ;

[0130] Table 1 presents the temporal tracking data of an example primary perturbation region. The area change rate sequence [0.15, 0.22, 0.18, 0.25] was smoothed and calculated, with an average of approximately 0.20 and a variance of approximately 0.0017. The overall decreasing slope of the Jaccard similarity coefficient sequence [0.850, 0.720, 0.650, 0.530] is approximately -0.105. The cumulative drift of the positional drift distance sequence [2.1, 1.8, 3.0, 2.5] is approximately 9.4 pixels. The temporal evolution feature vector is derived from these values. It can be understood that the continuous positive area growth is consistent with flame diffusion characteristics. The continuous decrease in the Jaccard similarity coefficient indicates a continuous change in shape, consistent with flame flickering characteristics. The cumulative drift indicates movement at the region center. These dynamic features help distinguish flames from static or regularly moving artificial light sources.

[0131] See Figure 3 The fire score curve in the figure is calculated by fusing the spatial feature vector and temporal evolution feature vector of the primary disturbance area. A quantitative score is generated through weighted voting of historical fire cases and nonlinear correction. The warning threshold and fire judgment threshold constitute the basis for the system's hierarchical decision-making. A higher score indicates that the area more closely matches the characteristics of a real fire, such as dynamic flame spread, irregular swaying, and positional drift. The curve trend fully reflects the dynamic process of a fire from its inception and development to stable combustion. The continuous rise in score and its breaking through the threshold are highly consistent with the physical evolution characteristics of flames. The fluctuation range reflects the natural disturbance characteristics of flame combustion, effectively distinguishing fires from interference factors such as regularly moving artificial light sources. This schematic diagram not only intuitively visualizes the hierarchical decision-making for warning and judgment, accurately supports the positioning mapping from image coordinates to geographic coordinates and the output of structured control commands, but also provides an intuitive basis for the performance verification of spatiotemporal feature extraction algorithms and temporal tracking models. It has significant engineering application value in improving detection reliability, reducing false alarm rates, assisting emergency decision-making, and system iterative optimization.

[0132] In one embodiment of the present invention, the process of calculating the fire score by integrating spatial and temporal domain analysis results is as follows: For each primary disturbance region, the spatial feature vector and the temporal evolution feature vector of the primary disturbance region are concatenated to form a combined feature vector. The dimension of the combined feature vector is the sum of the dimensions of the spatial feature vector and the temporal evolution feature vector. Several historical cases that are closest to the current combined feature vector in feature space distance are retrieved from the historical fire case database. The number of retrieval cases is set to 5, and the feature space distance is calculated using Euclidean distance. The final fire judgment result label corresponding to the historical case is read. The final fire judgment result label is either "fire" or "no fire". Based on the distance between the historical case and the current combined feature vector, different weights are assigned to each historical case, with closer cases having higher weights. A weighted vote is performed on the final fire judgment result labels of the historical cases, and the ratio of the total weight of the vote result "fire" to the total weight of all votes is calculated to obtain the primary fire probability based on case matching. The comprehensive disturbance intensity of the current primary disturbance region, the internal texture complexity features of the current primary disturbance region, and the average value of the area change rate sequence of the current primary disturbance region are nonlinearly combined and mapped to a correction factor. This correction factor is then used to adjust the primary fire probability based on case matching, resulting in a corrected fire score.

[0133] In practice, all primary disturbance areas with fire scores exceeding a preset fire judgment threshold (0.75) are selected. For each primary disturbance area meeting the threshold, a set of vertex coordinates of the convex polygon contour is extracted from the spatial feature vector of the primary disturbance area. The vertex coordinates are coordinates in the image pixel coordinate system. The vertex coordinate set is mapped from the image coordinate system to the actual geographic coordinate system through a pre-calibrated homography transformation matrix. The homography transformation matrix is ​​obtained through camera calibration, yielding the polygon boundary coordinates of the fire area in geographic space. For the same suspected fire source, if there are primary disturbance areas meeting the threshold in multiple consecutive frames, these primary disturbance areas in multiple consecutive frames are considered as different time observations of the same fire event. The geographic polygon boundary coordinates calculated in different frames of the fire event are fused, and the union of the polygon coordinates is taken as the final geographic location of the fire event. A fire event identifier is generated for each independent final geographic location; the fire event identifier is a globally unique string. The fire event identifier, the boundary coordinates of the final geographical location, the timestamp of the first detection of the fire event, and the average fire score corresponding to the fire event are bound together into a complete fire determination and location record. The fire determination and location record is encapsulated in a structured data format, using JSON format, and serves as the core content output for control commands.

[0134] In some embodiments, the weight allocation in weighted voting uses the reciprocal of the distance. The formula for calculating the nonlinear combination of the correction factor is:

[0135]

[0136] in: Indicates the correction factor. This indicates the overall disturbance intensity of the current primary disturbance region. This represents the internal texture complexity feature of the current primary perturbation region. This represents the average value of the area change rate sequence of the current primary disturbance region. , , It is the combination coefficient, the combination coefficient Set to 2.0, combination coefficient Set to 1.5, combination coefficient Set to 3.0, It is a natural constant. The final fire severity score. Through formula Calculation, where This is a primary fire probability based on case matching. The homography transformation matrix is ​​a 3x3 matrix. The fusion of geographic polygon boundary coordinates is achieved by calculating the minimum hull convex polygon of multiple polygon contour points. The structured data format of fire determination and location records includes event ID, a list of polygon vertex latitude and longitude, alarm time, and confidence score fields. In specific implementation, an example of a primary disturbance area is as follows: the distance, label, and assigned weight of the combined feature vector with the 5 most recent cases in the historical database are shown in Table 2.

[0137] Table 2: Distance, Label, and Weighting of Recent Cases

[0138] ;

[0139] Based on Table 2, a weighted vote was conducted. The total weight for the vote "It is a fire" was 0.45 + 0.30 + 0.05 = 0.80, and the total weight of all votes was 1.00. This was used to determine the initial fire probability based on case matching. Assuming the overall disturbance intensity in this region... Internal texture complexity features The average value of the area change rate series Calculate the correction factor. Then the fire situation score The score exceeded the fire severity threshold of 0.75. The vertex coordinates of the convex polygon outline were mapped using a homography transformation matrix to obtain a set of geographic coordinates. Since the fire source was detected in the subsequent five consecutive frames, the union of the polygon coordinates from these five frames was used to form the final geographic location area, and a record was generated as output. It can be understood that the combined feature vector comprehensively utilizes both spatial static features and temporal dynamic features. Historical case matching provides experience-based prior probabilities. Correction factors introduce intensity, texture, and growth features specific to the current region, providing personalized correction to the prior probabilities. The fire severity score is a normalized probability value. Coordinate mapping realizes the transformation from image perception to geographic information. Multi-frame observation fusion improves the robustness and accuracy of the location. Structured output facilitates integration with other systems.

[0140] See Figure 4 This paper presents the distribution characteristics and quantitative relationships of two core variables—feature spatial distance and assigned weights—in a visual recognition-based fire detection system. In the algorithm logic for fire case matching, feature spatial distance characterizes the similarity between the combined feature vector of the current primary disturbance area and similar cases in the historical fire case database. Its distribution range reflects the feature dispersion between different matched cases: the upper edge of the box reaches 4.0, the lower edge is approximately 1.2, the median is approximately 2.5, and the interquartile range covers the range of 1.8 to 3.1, reflecting the large span of spatiotemporal feature differences between different fire cases in actual scenarios. The algorithm needs to achieve accurate matching through multi-dimensional feature fusion. As the core parameter for weighted voting in case matching, the assigned weights are concentrated in the low value range of 0 to 0.5, with a median of approximately 0.2 and an interquartile range concentrated in the range of 0.1 to 0.3. This demonstrates that the algorithm strictly follows the allocation rule of "the closer the distance, the greater the weight," assigning high weights to nearby cases and low weights to distant cases. Furthermore, the concentration of the weight distribution further ensures the robustness of the weighted voting results. The comparison of the distribution of the two types of indicators clearly confirms the core mechanism of "case matching weighted probability + nonlinear correction of feature parameters" in the fire score calculation. The wide distribution of feature space distance adapts to the multi-case retrieval of complex fire scenarios, while the concentrated distribution of weight allocation ensures the effective weighting of retrieval results in the final score, providing reliable data support for the accurate triggering of the fire judgment threshold (0.75).

[0141] In one embodiment of the invention, the self-optimization phase during the continuous learning and initialization phase after the control command is issued, and the feedback learning phase, are executed after the control command triggers an external response, which refers to activating the on-site alarm or notifying the fire protection platform. A predetermined length of subsequent monitoring video segments is continuously collected, defined as 5 minutes of video data after the alarm. In these subsequent monitoring video segments, the actual development of the fire is confirmed manually or through more reliable independent sensors, such as on-site deployed thermal imaging fire detectors or smoke sensors. A fire truth label is generated, containing information on "real fire," "false alarm," and the actual spread range. From the beginning of the subsequent monitoring video segments, the actual evolution data of the areas in the primary disturbance area list within this time period is extracted. This actual evolution data includes actual spatial feature changes and temporal evolution trajectories. The actual evolution data is compared with the initially generated spatial feature vector and temporal evolution feature vector to calculate the feature prediction deviation. The feature prediction deviation, the initial fire score, and the fire truth label are correlated to form a feedback learning sample. Feedback learning samples are stored in a historical fire case database and used to update the weight allocation rules or nonlinear combination function parameters based on case matching in the fire scoring calculation model.

[0142] In practical implementation, the scene adaptive modeling phase is executed during system initialization. In a fire-free monitoring scenario, video data from a complete natural day is collected as a scene modeling sample, covering daytime, nighttime, and different lighting periods. The scene modeling sample is analyzed to identify periodically occurring interference sources in the scene, including vehicle lights, on / off lighting fixtures, and moving sunlight spots. The visual feature patterns and their temporal patterns of each periodic interference source are extracted. These visual feature patterns include the color, texture, spectral response, and motion characteristics of the interference source in the basic and auxiliary visual flows, generating an interference source feature-time pattern library. In subsequent real-time detection, when the spatial characteristics and occurrence time of a primary disturbance region identified in the basic and auxiliary visual flows match a pattern in the interference source feature-time pattern library with a pre-set interference matching threshold (set to 0.8), the overall disturbance intensity of the primary disturbance region is significantly attenuated, or it is directly removed from the primary disturbance region list. In the scene modeling samples, the natural fluctuation range of the steady-state feature model parameters under different weather conditions, including sunny, cloudy, and rainy days, is statistically analyzed. The steady-state feature model parameters refer to the mean and variance of the Gaussian components of the steady-state features. The natural fluctuation range is used as a dynamic threshold for updating the steady-state feature model. In real-time detection, a slow update of the steady-state feature model is triggered only when the color change of a pixel exceeds the natural fluctuation range, thus avoiding the model being contaminated by brief disturbances.

[0143] In some embodiments, calculating the feature prediction bias involves comparing multiple feature dimensions, and the formula for calculating the feature prediction bias is as follows:

[0144]

[0145] in: Indicates the average feature prediction bias. Indicates the number of features involved in the comparison. This represents the first element in the initially generated spatial feature vector or temporal evolution feature vector. 1 eigenvalue, This indicates the actual evolution data extracted from subsequent monitoring video segments corresponding to the [number]th [segment / section]. Each feature value is used. After associating the ground truth fire label with feature prediction bias and the initial fire score, the feedback learning samples are stored in a structured record format. Updating the historical fire case database includes adding new sample records and recalculating the distance-weight mapping relationship in the case-matching weight allocation rule. The nonlinear combination function parameters are updated using gradient descent to reduce feature prediction bias and the difference between the fire score and the ground truth label.

[0146] See Figure 5The figure presents the changing patterns of core indicators during the model's iterative optimization process, providing quantitative support for the effectiveness of feedback learning. Specifically, in the feedback learning phase, the system uses gradient descent to iteratively optimize the parameters of the nonlinear combination function in the fire scoring calculation model. The goal is to minimize the feature prediction bias, the difference between the fire score and the true fire label, and simultaneously improve the fire identification accuracy. The figure includes two core performance curves: the model loss curve corresponds to the loss function value during the gradient descent optimization process, which physically represents the comprehensive deviation between the model's prediction results (fire score, feature evolution prediction) and the true fire label and actual evolution data. As the curve trend shows, as the number of optimization iterations increases from 1 to 10, the model loss value continuously decreases from the initial approximately 0.176, eventually converging to approximately 0.058, exhibiting typical gradient descent convergence characteristics. This indicates that the model continuously corrects parameter biases during iteration, continuously improving prediction accuracy and ultimately reaching a stable low-loss state. The identification accuracy curve corresponds to the model's classification accuracy for real fires and false alarm scenarios during the iterative optimization process. As the number of iterations increased, the recognition accuracy gradually rose from an initial approximately 0.73, eventually stabilizing at approximately 0.92. This showed a completely corresponding negative correlation with the decrease in the model loss value, validating the effectiveness of gradient descent optimization: by continuously reducing the model loss, the system's fire recognition accuracy was significantly improved, ultimately achieving highly reliable fire assessment performance. In the system parameter configuration, the number of iterations for gradient descent optimization was set to 10, and an adaptive decay strategy was adopted for the learning rate. The initial learning rate was set to 0.01, decaying by 50% every two iterations to ensure rapid convergence of the model in the early stages of iteration and stable fine-tuning in the later stages, avoiding overfitting. This performance curve validates the core value of the feedback learning stage: by using real fire evolution data and ground truth labels as feedback, the model parameters are continuously optimized, achieving a closed loop of continuous self-evolution from initial detection, significantly improving the robustness and accuracy of fire detection, and effectively reducing the false alarm rate.

[0147] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A visual recognition-based intelligent fire detection and positioning method, characterized in that, Including the following steps: Video source signals are generated from continuous frame images containing panoramic views of the scene acquired by a video sensor array. The basic visual stream for preliminary analysis and the auxiliary visual stream for cross-validation are separated from the video source signal. The basic visual flow is scanned frame by frame, and a steady-state feature model of the scene is generated through dynamic background modeling. Perform spectral transformation on the auxiliary visual stream to generate an enhanced spectral response stream; Features matching the preset potential fire spectral range are extracted from the enhanced spectral response stream to generate a potential fire spectral feature set. The potential fire spectral feature set is then compared point by point with the steady-state feature model to mark the perturbation regions that exceed the steady-state features and generate a primary perturbation region list. Based on the list of primary disturbance regions, a multi-stage analysis of the video source signal is initiated, which includes a spatial domain analysis stage and a temporal domain analysis stage. In the spatial domain analysis phase, the boundary contour, internal texture complexity, and color statistical distribution of each primary perturbation region are calculated. In the temporal analysis phase, the area change, shape evolution path, and positional drift trajectory of each primary perturbation region are tracked across multiple frames of images. Based on the combined spatial domain analysis results and temporal domain analysis results, the fire score for each primary disturbance area is calculated. Based on the spatiotemporal information of the primary disturbance area where the fire score exceeds the decision threshold, the fire determination result and geographic location coordinates are generated, and control commands are output. The step of scanning the basic visual stream frame by frame and generating a steady-state feature model of the scene through dynamic background modeling includes: Starting from the beginning of the basic visual flow, select a series of consecutive frames as reference frames for background modeling; Extract the color value sequence of each pixel in the background modeling reference frame across consecutive frames; For each pixel, a Gaussian mixture model is fitted to its color value sequence to generate a set of multimodal steady-state Gaussian components for each pixel. The set of multimodal steady-state Gaussian components of a pixel is weighted and evaluated, and the steady-state Gaussian component with the highest weight is defined as the steady-state feature Gaussian component of the pixel. Collect the Gaussian components of the steady-state features of all pixels to form a two-dimensional steady-state feature matrix; Calculate the similarity of the Gaussian components of the steady-state features of adjacent pixels in the color space in the two-dimensional steady-state feature matrix, and construct the boundary of the steady-state feature region based on the similarity; Based on the boundaries of the steady-state feature regions, adjacent pixels with similar steady-state Gaussian components are grouped into steady-state feature blocks; Each steady-state feature block is assigned a unique feature block identifier, and the average color value and spatial covariance matrix corresponding to the feature block identifier are recorded, which together form the steady-state feature model.

2. The intelligent fire detection and location method based on visual recognition according to claim 1, characterized in that, The step of performing spectral transformation processing on the auxiliary visual stream to generate an enhanced spectral response stream includes: Infrared spectral images and short-wave infrared spectral images are separated from the raw auxiliary visual stream acquired by the multispectral camera. Histogram equalization is performed on the infrared spectral band image to enhance the overall image contrast and generate the first processed spectral image. Adaptive gamma correction is applied to shortwave infrared spectral images to highlight the energy characteristics of preset bands and generate second-processed spectral images. The first processed spectral image and the second processed spectral image are then fused at the pixel level to generate a fused spectral image; A filter based on a preset fire spectral reference is applied to the fused spectral image, the preset fire spectral reference being generated based on a standard flame spectral database; Calculate the response intensity of each pixel in the fused spectral image with the preset fire spectral reference to obtain a preliminary spectral response map; Multi-scale spatial morphological operations are performed on the preliminary spectral response map to remove isolated noise response points and connect adjacent strong response regions to form connected regions. Extract the geometric center coordinates, bounding rectangle, and response intensity peak of each connected region to form a list of spectral feature descriptors; Based on the list of spectral feature descriptors, the corresponding response regions are marked on each frame of the original auxiliary visual stream to generate the enhanced spectral response stream.

3. The intelligent fire detection and location method based on visual recognition according to claim 2, characterized in that, The process of generating a potential fire spectral feature set and comparing it point-by-point with the steady-state feature model to mark perturbation regions exceeding the steady-state features includes: Iterate through each response region defined by the list of spectral feature descriptors in the enhanced spectral response stream; Locate the pixel position in the basic visual flow that corresponds to the geometric center coordinates of the response region; From the steady-state feature model, read the mean and variance of the steady-state feature Gaussian components corresponding to the pixel position; Read the actual color value of the pixel location from the current frame of the basic visual stream; Calculate the probability that the actual color value falls within the color distribution range defined by the corresponding steady-state Gaussian component, and generate a steady-state conformity value; If the steady-state compliance value is lower than the set steady-state deviation threshold, then the pixel position is determined to be a steady-state deviation point; Collect all steady-state deviations and mark the response regions to which they belong as suspicious regions; For all suspicious regions in the basic visual flow, the steady-state conformity value is repeatedly calculated, and the number of steady-state deviation points and the proportion of the total area in each suspicious region are counted. The overall perturbation intensity of the suspected region is obtained by weighted calculation by combining the peak response intensity of the suspected region in the enhanced spectral response stream. Suspicious regions whose overall disturbance intensity exceeds the disturbance threshold are filtered to form the primary disturbance region list. The entries in the primary disturbance region list include the boundary information and overall disturbance intensity of each region.

4. The intelligent fire detection and location method based on visual recognition according to claim 3, characterized in that, The calculation of the boundary contour, internal texture complexity, and color statistical distribution of each primary perturbation region includes: For each primary perturbation region, extract the set of pixels from its boundary information; Perform convex hull calculation on the set of pixels to generate a convex polygonal contour of the primary perturbation region, and record the coordinates of each vertex of the convex polygonal contour; Based on the convex polygon profile, the ratio of its area to its perimeter is calculated as a profile compactness feature. Within the boundaries of the primary perturbation region, extract grayscale image sub-blocks; The local binary pattern operator is applied to the image sub-blocks, the histograms of different local binary pattern codes are statistically analyzed, and the entropy value of the histograms is calculated as an internal texture complexity feature. Within the boundary of the primary perturbation region, the mean, variance, skewness, and kurtosis of all pixels in multiple color channels are statistically analyzed to form a color statistical distribution vector. The contour compactness feature, internal texture complexity feature, and color statistical distribution vector are combined to generate the spatial feature vector of the primary perturbation region.

5. The intelligent fire detection and location method based on visual recognition according to claim 4, characterized in that, The tracking of the area change, shape evolution path, and positional drift trajectory of each primary perturbation region across multiple frames of images includes: In the subsequent frames of the base visual stream containing the primary perturbation region, each primary perturbation region is dynamically tracked. In the current frame, starting from the center position of the primary perturbation region determined in the previous frame, template matching is performed within a preset neighborhood range to determine the approximate position of the primary perturbation region in the current frame. At the approximate location, the color statistical distribution vector of the previous frame of the primary perturbation region is used as a reference to perform region growing segmentation to obtain the segmented region of the current frame. Calculate the area of ​​the segmented region in the current frame and compare it with the area in the previous frame to calculate the rate of change of area. Extract the convex polygon contour of the segmented region in the current frame, calculate its Jaccard similarity coefficient with the convex polygon contours of the previous several frames, and describe the shape evolution path. Record the geometric center coordinates of the segmented region in the current frame, calculate the Euclidean distance between it and the geometric center coordinates of the previous frame, and describe the inter-frame position drift. Smoothing filters are applied to the area change rate sequence, Jaccard similarity coefficient sequence, and position drift distance sequence of consecutive frames to remove noise points; Calculate the mean and variance of the smoothed area change rate sequence, calculate the overall descent slope of the Jaccard similarity coefficient sequence, and calculate the cumulative drift of the position drift distance sequence. The mean and variance of the area change rate sequence, the overall descent slope of the Jaccard similarity coefficient sequence, and the cumulative drift of the position drift distance sequence are combined to generate the temporal evolution feature vector of the primary perturbation region.

6. The intelligent fire detection and location method based on visual recognition according to claim 5, characterized in that, The combined spatial domain analysis results and temporal domain analysis results are used to calculate the fire score for each primary disturbance area, including: For each primary perturbation region, its spatial feature vector and temporal evolution feature vector are concatenated to form a combined feature vector; Retrieve from the historical fire case database several historical cases that are closest in feature space distance to the current combined feature vector; Read the final fire situation assessment result tags corresponding to historical cases; Based on the distance between historical cases and the current combined feature vector, each historical case is assigned a different weight, with the weight increasing as the distance increases. The final fire situation determination labels of historical cases are weighted and voted on. The ratio of the total weight of the vote result "is a fire" to the total weight of all votes is calculated to obtain the initial fire situation probability based on case matching. The average value of the comprehensive disturbance intensity, internal texture complexity features, and area change rate sequence of the current primary disturbance region is nonlinearly combined and mapped into a correction factor. The correction factor is used to adjust the initial fire probability based on case matching to obtain the corrected fire score.

7. The intelligent fire detection and location method based on visual recognition according to claim 6, characterized in that, The generation of fire determination results and geographic location coordinates includes: All primary disturbance areas whose fire scores exceed the preset fire judgment threshold are selected; For each primary perturbation region that meets the threshold condition, extract the set of vertex coordinates of the convex polygon contour from its spatial feature vector. The vertex coordinate set is mapped from the image coordinate system to the actual geographic coordinate system through a pre-calibrated homography transformation matrix to obtain the polygon boundary coordinates of the fire area in geographic space. For the same suspected fire source, if there are primary disturbance regions that meet the threshold conditions in multiple consecutive frames, they are considered as different time observations of the same fire event. The geographical polygon boundary coordinates calculated in different frames of the fire event are fused together, and the union of the polygon coordinates is taken as the final geographical location of the fire event. Generate a fire event identifier for each individual final location geographic area; The fire event identifier, the boundary coordinates of the final location geographic area, the timestamp of the first detection, and the corresponding average fire score are bound together to form a complete fire determination and location record; The fire detection and location records are encapsulated into a structured data format and output as the core content of the control commands.

8. The intelligent fire detection and location method based on visual recognition according to claim 7, characterized in that, It also includes a feedback learning phase after the output control command, the feedback learning phase including: After the control command triggers an external response, it continuously collects subsequent monitoring video segments of a predetermined length; In subsequent monitoring video segments, the actual development of the fire is confirmed manually or through more reliable independent sensors, including whether the fire is real and the actual extent of its spread, and a fire truth value label is generated. From the beginning of subsequent monitoring video segments, extract the actual evolution data of the regions in the list of primary disturbance regions within this time period, including actual spatial feature changes and temporal evolution trajectories; The actual evolution data is compared with the initially generated spatial feature vector and temporal evolution feature vector to calculate the feature prediction bias; The feature prediction bias, the initial fire score, and the fire truth label are correlated to form a feedback learning sample. Feedback learning samples are stored in a historical fire case database and used to update the weight allocation rules or nonlinear combination function parameters based on case matching in the fire scoring calculation model.

9. The intelligent fire detection and location method based on visual recognition according to claim 8, characterized in that, It also includes a scene adaptive modeling phase executed during system initialization, which includes: In a fire-free monitoring scenario, video data from a complete natural day is collected as a sample for scene modeling. Analyze the scene modeling samples to identify periodically occurring interference sources in the scene, including vehicle lights, the on of lighting fixtures, and moving sunlight spots; Extract the visual feature patterns of each periodic interference source and their temporal occurrence patterns to generate an interference source feature-time pattern library; In subsequent real-time detection, when the spatial characteristics and occurrence time of the primary disturbance region identified in the basic visual flow and the auxiliary visual flow match a certain pattern in the interference source feature-time pattern library with a preset interference matching threshold, the comprehensive disturbance intensity of the primary disturbance region will be significantly reduced, or it will be directly removed from the list of primary disturbance regions. In the scene modeling samples, the natural fluctuation range of steady-state characteristic model parameters under different weather conditions was statistically analyzed; The natural fluctuation range is used as the dynamic threshold for updating the steady-state feature model. In real-time detection, the steady-state feature model is only slowly updated when the color change of a pixel exceeds this dynamic threshold, thus avoiding the model being contaminated by brief interference.

Citation Information

Patent Citations

  • Linear fire detection method and system based on infrared spectrum analysis and infrared spectrum detector

    CN120431673A

  • Fire prevention and control method and system using computer vision

    CN121147856A