A fire safety inspection method and system based on machine vision recognition
Through the fire safety inspection method based on machine vision recognition, a multi-spectral camera and sensor network are used to conduct comprehensive perception and feature analysis, combined with multi-modal feature fusion and situational awareness warning, the problems of insufficient early fire hazard recognition capabilities and high false alarm rates in the existing technology are solved, and more efficient and reliable fire detection and early warning are achieved.
Patent Information
- Application Number
- CN202510090021.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-01-21
AI Technical Summary
The existing fire safety inspection methods are difficult to effectively capture the subtle characteristics of early fire hazards and reduce unnecessary alarms in the early stages.
The fire safety inspection method based on machine vision recognition is adopted to conduct comprehensive perception through multi-spectral cameras and sensor networks, and multi-spectral dynamic perception, motion amplification technology and thermal texture feature analysis are used, combining multi-modal feature fusion and situational awareness warning to generate intelligent early warning reports.
It significantly improves the efficiency and reliability of early detection and early warning of fires, and can identify fire hazards earlier and more accurately and reduce false alarms.
Smart Images

Figure CN119992032B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fire safety inspection, and particularly to a fire safety inspection method and system based on machine vision recognition. Background Art
[0002] Traditional fire safety inspections mainly rely on inspectors to conduct visual, auditory, olfactory and other sensory inspections in the inspection area regularly or irregularly, to check whether fire-fighting facilities are intact, whether there are any violations, whether there are abnormal odors or smoke, etc. This method is inefficient, highly subjective, prone to fatigue and negligence, difficult to cover all areas and times, and it is difficult to detect early and subtle fire hazards. Currently, various types of fire detectors, such as smoke detectors, heat detectors, flame detectors, etc., are also used to form an automatic fire alarm system. The response speed is faster than manual inspections and can achieve 24-hour monitoring.
[0003] However, the existing methods have problems of insufficient early fire hazard recognition ability and high false alarm rate:
[0004] 1. Insufficient ability to recognize early fire hazards: Existing fire detection methods mainly focus on obvious smoke, flames and high temperatures, while the characteristics of early fires are often very subtle, such as slight temperature rise, heat disturbance, small amounts of smoke or the release of specific gases. These subtle characteristics are difficult to be captured by traditional methods, resulting in early fires being difficult to be detected in time.
[0005] 2. High false alarm rate: Existing fire detection methods are easily affected by environmental interference, such as light changes, background movement, other high-temperature objects, etc., resulting in false alarms. For example, misjudging vehicle exhaust, sunlight reflection, electric welding sparks, etc. as fires. Summary of the Invention
[0006] Based on this, it is necessary to provide a fire safety inspection method and system based on machine vision recognition to solve at least one of the above technical problems.
[0007] To achieve the above object, a fire safety inspection method based on machine vision recognition includes the following steps:
[0008] Step S1: Deploy inspection equipment in the inspection area to obtain a camera deployment plan and a sensor network deployment plan; perform multi-spectral dynamic perception according to the camera deployment plan to obtain a temporal spectral stream;
[0009] Step S2: Perform multi-spectral channel motion estimation on the temporal spectral stream, and suppress the motion vector noise to obtain a filtered motion vector map; perform motion enhancement based on spectral characteristics on the filtered motion vector map to obtain a spectrally enhanced motion vector; apply the motion magnification algorithm to the temporal spectral stream according to the spectrally enhanced motion vector, and construct an enhanced motion vector field to obtain an enhanced motion vector field;
[0010] Step S3: Extract the thermal texture features from the temporal spectral stream to obtain a thermal texture feature map; perform dynamic change analysis on the thermal texture feature map to obtain a thermal texture change map; perform thermal texture anomaly localization and characterization on the preprocessed thermal image sequence according to the thermal texture change map to obtain a thermal texture anomaly map;
[0011] Step S4: Perform multi-modal feature fusion on the temporal spectral stream, the enhanced motion vector field, and the thermal texture anomaly map to obtain a fused feature vector; perform risk confidence prediction and generation on the fused feature vector to obtain a risk confidence heat map;
[0012] Step S5: Collect and preprocess the context information of the inspection area according to the sensor network deployment plan to obtain a context information dataset; extract the risk area information from the risk confidence heat map to obtain associated risk area information; perform context-aware early warning according to the associated risk area information and the context information dataset to obtain an intelligent early warning report.
[0013] Through the deployment of multi-spectral cameras and sensor networks, the present invention realizes the comprehensive perception and information collection of the inspection area. The multi-spectral dynamic perception and motion magnification technology effectively highlights the weak motion information in the early stage of a fire. The extraction of thermal texture features and dynamic change analysis capture the thermal anomaly features related to the fire. The multi-modal feature fusion technology combines spectral, motion, and thermal texture information, improving the accuracy of fire risk identification. The collection and preprocessing of context information provide richer context information for risk assessment, and through context-aware early warning, the intelligent assessment and early warning of fire risks are realized. Finally, an intelligent early warning report including the location of the risk area, the early warning level, the analysis of potential causes, and disposal suggestions is generated, thus significantly improving the efficiency and reliability of early fire detection and early warning. Therefore, the present invention provides a fire safety inspection method based on machine vision recognition, which effectively solves the problems of insufficient ability to identify early fire hazards and high false alarm rates in existing methods through multi-modal information fusion, context awareness, and the capture of early subtle features, realizing earlier and more accurate identification and early warning of fire hazards.
[0014] Preferably, step S1 includes the following steps:
[0015] Step S11: Deploy multi-spectral cameras in the inspection area and deploy a sensor network to obtain a camera deployment plan and a sensor network deployment plan;
[0016] Step S12: Rapidly acquire a multi-spectral image sequence according to the camera deployment plan to obtain an original multi-spectral image stream;
[0017] Step S13: Preprocess the original multi-spectral image stream to obtain a corrected multi-spectral image sequence;
[0018] Step S14: Align and fuse the channels of the corrected multi-spectral image sequence to obtain a multi-spectral image stack;
[0019] Step S15: Integrate the temporal spectral information of the multi-spectral image stack to obtain a temporal spectral stream.
[0020] Through the multi-spectral camera deployment and sensor network deployment plan, the present invention realizes the full coverage and precise monitoring of the inspection area, and effectively captures early fire signs. Rapidly acquiring the multi-spectral image sequence and performing preprocessing, including radiometric calibration, denoising, and geometric correction, ensures the accuracy and reliability of the image data. The alignment and fusion of multi-spectral image channels, as well as the integration of temporal spectral information, construct a multi-dimensional dataset containing time, space, and spectral information, providing a rich data basis for subsequent early fire recognition and warning. Through precise time synchronization and lossless compression storage, the original information is maximally retained, improving the accuracy and efficiency of early fire recognition.
[0021] Preferably, step S2 includes the following steps:
[0022] Step S21: Perform multi-spectral channel motion estimation on the temporal spectral stream to obtain a channel motion vector map;
[0023] Step S22: Suppress and screen the motion vector noise of the channel motion vector map to obtain a filtered motion vector map;
[0024] Step S23: Perform motion enhancement based on spectral characteristics on the filtered motion vector map according to the temporal spectral stream to obtain a spectrally enhanced motion vector;
[0025] Step S24: Use the spectrally enhanced motion vector as guiding information to apply a motion magnification algorithm to the temporal spectral stream to obtain an amplified image sequence;
[0026] Step S25: Construct an enhanced motion vector field for the amplified image sequence according to the spectrally enhanced motion vector to obtain an enhanced motion vector field.
[0027] Through multi - spectral channel motion estimation, noise suppression and screening, as well as motion enhancement based on spectral characteristics, the present invention effectively extracts the weak motion information caused by early - stage fires and suppresses the interference of background noise and irrelevant motions. By using the spectral - enhanced motion vector to guide the motion magnification algorithm, the tiny motions related to fires are highlighted, making them more visually obvious for easy observation and recognition. The finally constructed enhanced motion vector field integrates spectral information and motion information, providing a more accurate basis for subsequent fire feature extraction and recognition, and improving the sensitivity and reliability of early - stage fire detection.
[0028] Preferably, step S23 includes the following steps:
[0029] Step S231: Calculate the spectral feature response map for the temporal spectral flow to obtain a set of spectral feature response maps;
[0030] Step S232: Construct a spectral similarity weight map based on the set of spectral feature response maps to obtain a set of spectral similarity weight maps;
[0031] Step S233: Perform the fusion of spectral weights and motion vectors on the filtered motion vector map according to the set of spectral similarity weight maps to obtain a spectrally weighted motion vector map;
[0032] Step S234: Perform temporal motion vector smoothing and accumulation on the spectrally weighted motion vector map to obtain a temporally smoothed motion vector map;
[0033] Step S235: Perform motion vector enhancement based on the temporally smoothed motion vector map and the filtered motion vector map to obtain spectrally enhanced motion vectors.
[0034] By calculating the spectral feature response map and constructing the spectral similarity weight map, the present invention integrates spectral information into the calculation of motion vectors, effectively enhancing the motion information related to fire characteristics. Through the weighted - average fusion strategy, the spectral weights of various fire characteristics are combined, improving the accuracy of motion vectors. Temporal motion vector smoothing and accumulation effectively suppress noise interference using Kalman filtering and highlight the persistent motion information. The final motion vector enhancement step further magnifies the motions related to fires while retaining the original motion information, providing a more reliable basis for subsequent motion magnification and fire recognition, thereby improving the sensitivity and robustness of fire detection.
[0035] Preferably, step S3 includes the following steps:
[0036] Step S31: Extract the infrared image sequence from the temporal spectral flow and perform inverse geometric transformation using the enhanced motion vector field to obtain the aligned infrared image sequence;
[0037] Step S32: Perform thermal image preprocessing on the aligned infrared image sequence to obtain the preprocessed thermal image sequence;
[0038] Step S33: Extract thermal texture features from the preprocessed thermal image sequence to obtain a thermal texture feature map;
[0039] Step S34: Perform dynamic change analysis of the thermal texture on the thermal texture feature map to obtain a thermal texture change map;
[0040] Step S35: Locate and characterize thermal texture anomalies in the preprocessed thermal image sequence according to the thermal texture change map to obtain a thermal texture anomaly map.
[0041] In the present invention, by extracting infrared images from the temporal spectral flow and performing inverse geometric transformation, the image distortion caused by motion magnification is effectively eliminated, and the thermal information related to fire is extracted. Thermal image preprocessing enhances the image contrast and lays a foundation for subsequent feature extraction. The LBP operator is used to extract thermal texture features and analyze their dynamic changes, effectively capturing the thermal texture anomalies in the early stage of fire. Finally, through thermal texture anomaly location and characterization, combined with temperature and texture features, the accurate identification and classification of potential fire areas are realized, improving the accuracy and reliability of early fire detection.
[0042] Preferably, step S34 includes the following steps:
[0043] Step S341: Calculate a texture difference feature map for the thermal texture feature map to obtain a texture difference feature map;
[0044] Step S342: Establish a motion amplitude feature map for the enhanced motion vector field to obtain a motion amplitude feature map;
[0045] Step S343: Perform fusion of texture difference and motion amplitude on the texture difference feature map and the motion amplitude feature map to obtain a fused feature score map;
[0046] Step S344: Perform time series anomaly detection on the fused feature score map to obtain an anomaly score map;
[0047] Step S345: Generate a thermal texture change map according to the anomaly score map to obtain a thermal texture change map.
[0048] The present invention effectively combines thermal texture changes and motion information by calculating the texture difference feature map and the motion amplitude feature map and fusing them, thus more comprehensively depicting the dynamic characteristics in the early stage of a fire. By using the time series anomaly detection method to analyze the change trend of the fused feature scores, it effectively distinguishes normal thermal texture fluctuations from abnormal changes caused by a fire. The finally generated thermal texture change map highlights potential fire areas in the form of anomaly scores, providing a more reliable basis for early fire warning and improving the sensitivity and accuracy of fire detection.
[0049] Preferably, step S35 includes the following steps:
[0050] Step S351: Generate a binary anomaly mask for the thermal texture change map to obtain a binary anomaly mask;
[0051] Step S352: Perform component labeling connection on the binary anomaly mask to obtain a connected component labeling map;
[0052] Step S353: Extract abnormal region images from the preprocessed thermal image sequence according to the connected component labeling map to obtain a set of abnormal region image blocks;
[0053] Step S354: Calculate the regional temperature statistical features for the set of abnormal region image blocks to obtain a regional temperature feature table;
[0054] Step S355: Calculate the regional texture features for the set of abnormal region image blocks to obtain a regional texture feature table;
[0055] Step S356: Classify and label the abnormal patterns according to the regional temperature feature table and the regional texture feature table to obtain a labeled abnormal region map;
[0056] Step S357: Generate a thermal texture anomaly atlas according to the labeled abnormal region map, the regional temperature feature table, and the regional texture feature table to obtain a thermal texture anomaly atlas.
[0057] The present invention realizes the precise positioning, characterization, and classification of thermal texture anomalies through a series of steps. First, regions with significant thermal texture changes are extracted through binarization and connected component analysis. Then, the temperature statistical features and texture features of each abnormal region are calculated respectively, providing a basis for subsequent classification. Using a pre-trained classification model, the abnormal regions are classified and labeled, effectively distinguishing different types of early fire phenomena. The finally generated thermal texture anomaly atlas not only contains the location information of the abnormal regions but also provides detailed temperature and texture features, as well as abnormal type labels, providing more comprehensive and intuitive information for early fire warning and emergency response, thereby improving the efficiency and accuracy of fire detection.
[0058] Preferably, step S4 includes the following steps:
[0059] Step S41: Align the multi-modal feature data of the temporal spectral stream, enhanced motion vector field, and thermal texture anomaly map to obtain the aligned multi-modal feature set;
[0060] Step S42: Extract multi-modal deep features from the aligned multi-modal feature set to obtain multi-modal deep features;
[0061] Step S43: Perform feature-level fusion on the multi-modal deep features and conduct feature interaction learning to obtain the fused feature vector;
[0062] Step S44: Predict and generate the risk confidence level for the fused feature vector to obtain the risk confidence heat map.
[0063] Through the alignment of multi-modal feature data, the present invention ensures the spatial and temporal consistency of the temporal spectral stream, enhanced motion vector field, and thermal texture anomaly map, laying a foundation for subsequent multi-modal feature fusion. The deep learning model is used to extract the deep features of different modalities respectively, effectively capturing the fire-related information contained in different data sources. Based on the attention mechanism, feature-level fusion and interaction learning not only fuse the features of different modalities but also learn the relationships between them, highlighting important features and suppressing noise interference. The finally generated risk confidence heat map intuitively shows the spatial distribution of fire risks, providing a reliable decision-making basis for early fire warning and precise prevention and control, and significantly improving the efficiency and accuracy of fire detection.
[0064] Preferably, step S5 includes the following steps:
[0065] Step S51: Collect and preprocess the context information of the inspection area according to the sensor network deployment plan to obtain the context information data set;
[0066] Step S52: Extract the risk area from the risk confidence heat map to obtain the risk area data; match the spatial information of the risk area data and the context information data set, and conduct risk area association to obtain the associated risk area information;
[0067] Step S53: Perform context factor weighting according to the associated risk area information to obtain the context factor weighted data; perform risk correction according to the context factor weighted data to obtain the context-corrected risk confidence level;
[0068] Step S54: Divide the warning level according to the context-corrected risk confidence level to obtain the warning level division information; dynamically adjust the threshold of the warning level division information according to the context information data set to obtain the warning level information;
[0069] Step S55: Generate and push an intelligent early warning report based on the early warning level information, associated risk area information, and temporal spectral flow to obtain an intelligent early warning report.
[0070] The present invention realizes the intelligent assessment and early warning of fire risks by combining the context information collected by the sensor network, the risk confidence heat map, and the temporal spectral flow. The collection and preprocessing of context information provide rich context information for risk correction. The matching of spatial information and the association of risk areas link risks with specific scenarios and devices, achieving accurate risk positioning. The weighting of context factors and risk correction effectively improve the accuracy of risk assessment. The division of early warning levels and the dynamic adjustment of thresholds optimize the early warning levels according to real-time context information, avoiding false alarms and missed alarms. Finally, the generation and push of the intelligent early warning report provide timely, accurate, and comprehensive fire risk information and disposal suggestions for fire safety management personnel, thus effectively improving the efficiency and reliability of fire prevention and control.
[0071] Preferably, the present invention also provides a fire safety inspection system based on machine vision recognition for performing the fire safety inspection method based on machine vision recognition as described above. The fire safety inspection system based on machine vision recognition includes:
[0072] A multi-spectral dynamic perception module for deploying inspection devices in the inspection area to obtain a camera deployment plan and a sensor network deployment plan; performing multi-spectral dynamic perception according to the camera deployment plan to obtain a temporal spectral flow;
[0073] A micro-motion amplification module for performing multi-spectral channel motion estimation on the temporal spectral flow and suppressing motion vector noise to obtain a filtered motion vector map; performing motion enhancement based on spectral characteristics on the filtered motion vector map to obtain a spectrally enhanced motion vector; applying a motion amplification algorithm to the temporal spectral flow according to the spectrally enhanced motion vector and constructing an enhanced motion vector field to obtain an enhanced motion vector field;
[0074] A thermal texture anomaly analysis module for extracting thermal texture features from the temporal spectral flow to obtain a thermal texture feature map; performing dynamic change analysis of the thermal texture on the thermal texture feature map to obtain a thermal texture change map; performing thermal texture anomaly localization and characterization on the preprocessed thermal image sequence according to the thermal texture change map to obtain a thermal texture anomaly atlas;
[0075] A multi-modal feature fusion module for performing multi-modal feature fusion on the temporal spectral flow, the enhanced motion vector field, and the thermal texture anomaly atlas to obtain a fused feature vector; performing risk confidence prediction and generation on the fused feature vector to obtain a risk confidence heat map;
[0076] The situation awareness warning module is used to collect and preprocess situation information in the inspection area according to the sensor network deployment plan to obtain a situation information data set; extract risk area information from the risk confidence heat map to obtain associated risk area information; and perform situation awareness warning based on the associated risk area information and the situation information data set to obtain an intelligent warning report. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] Figure 1 It is a schematic flow chart of the steps of a fire safety inspection method based on machine vision recognition;
[0078] Figure 2 It is a schematic detailed implementation step flow chart of step S2 in the present invention;
[0079] Figure 3 It is a schematic detailed implementation step flow chart of step S3 in the present invention.
[0080] The realization of the object, functional features and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0081] The technical method of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative work belong to the scope of protection of the present invention.
[0082] In addition, the accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.
[0083] It should be understood that although the terms "first", "second", etc. may be used here to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit can be called the second unit, and similarly the second unit can be called the first unit. The term "and / or" used here includes any and all combinations of one or more of the listed related items.
[0084] To achieve the above object, please refer to Figures 1 to 3, a fire safety inspection method based on machine vision recognition, comprising the following steps:
[0085] Step S1: Deploy inspection equipment in the inspection area to obtain a camera deployment plan and a sensor network deployment plan; perform multi-spectral dynamic perception according to the camera deployment plan to obtain a time-series spectral stream;
[0086] Step S2: Perform multi-spectral channel motion estimation on the time-series spectral stream, and suppress motion vector noise to obtain a filtered motion vector map; perform motion enhancement based on spectral characteristics on the filtered motion vector map to obtain a spectrally enhanced motion vector; apply a motion magnification algorithm to the time-series spectral stream according to the spectrally enhanced motion vector, and construct an enhanced motion vector field to obtain an enhanced motion vector field;
[0087] Step S3: Extract thermal texture features from the time-series spectral stream to obtain a thermal texture feature map; perform dynamic change analysis on the thermal texture feature map to obtain a thermal texture change map; perform thermal texture anomaly localization and characterization on the preprocessed thermal image sequence according to the thermal texture change map to obtain a thermal texture anomaly map;
[0088] Step S4: Perform multi-modal feature fusion on the time-series spectral stream, the enhanced motion vector field, and the thermal texture anomaly map to obtain a fused feature vector; perform risk confidence prediction and generation on the fused feature vector to obtain a risk confidence heat map;
[0089] Step S5: Collect and preprocess context information in the inspection area according to the sensor network deployment plan to obtain a context information data set; extract risk area information from the risk confidence heat map to obtain associated risk area information; perform context-aware warning according to the associated risk area information and the context information data set to obtain an intelligent warning report.
[0090] In the embodiment of the present invention, referring to Figure 1 as shown, it is a schematic flow chart of the steps of the fire safety inspection method based on machine vision recognition of the present invention. In this example, the fire safety inspection method based on machine vision recognition comprises the following steps:
[0091] Step S1: Deploy inspection equipment in the inspection area to obtain a camera deployment plan and a sensor network deployment plan; perform multi-spectral dynamic perception according to the camera deployment plan to obtain a time-series spectral stream;
[0092] In the embodiments of the present invention, a multispectral camera with a specific wavelength band and resolution is selected, optimized and installed according to the layout of the inspection area and environmental factors, and initial calibration is carried out to form a camera deployment plan. A wireless network including multiple sensors is synchronously deployed to form a sensor network deployment plan. Subsequently, multispectral images are synchronously collected at a high frame rate, and preprocessing such as radiometric calibration, denoising and geometric correction is carried out, and channel alignment and fusion are carried out, and finally integrated into a temporal spectral stream including multispectral channel information.
[0093] Step S2: Perform multispectral channel motion estimation on the temporal spectral stream, and suppress motion vector noise to obtain a filtered motion vector map; perform motion enhancement based on spectral characteristics on the filtered motion vector map to obtain a spectrally enhanced motion vector; apply a motion magnification algorithm to the temporal spectral stream according to the spectrally enhanced motion vector, and construct an enhanced motion vector field to obtain an enhanced motion vector field;
[0094] In the embodiments of the present invention, high-precision motion estimation is performed on each spectral channel respectively, and methods such as median filtering and amplitude threshold are used to suppress noise and filter out small motion vectors. Then, combined with spectral feature information, the motion vectors related to the early fire characteristics are selectively enhanced. Next, the motion magnification algorithm is guided by the enhanced motion vectors to highlight the small motions. Finally, the motion of the magnified image is re-estimated and fused with the spectrally enhanced motion vectors to construct an enhanced motion vector field containing rich small motion information.
[0095] Step S3: Extract thermal texture features from the temporal spectral stream to obtain a thermal texture feature map; perform thermal texture dynamic change analysis on the thermal texture feature map to obtain a thermal texture change map; perform thermal texture anomaly localization and characterization on the preprocessed thermal image sequence according to the thermal texture change map to obtain a thermal texture anomaly map;
[0096] In the embodiments of the present invention, first, the infrared image sequence in the temporal spectral stream is extracted and preprocessed. Then, the local texture features of the thermal image are extracted by using the LBP operator, and the difference in texture features between adjacent frames is calculated to obtain a thermal texture change map. A threshold is set to generate a binary anomaly mask, and independent anomaly regions are identified through connected component analysis. Then, image patches of these anomaly regions are extracted from the preprocessed thermal image, and their average temperature, maximum temperature, temperature standard deviation and LBP histogram statistical features are calculated. Finally, based on these temperature and texture features, the anomaly patterns are classified by using a pre-trained classification model or set rules, such as classified into overheating, smoke or flame, and different types of anomaly regions are marked on the thermal image to generate a thermal texture anomaly map containing position, type and detailed feature information.
[0097] Step S4: Perform multi-modal feature fusion on the temporal spectral stream, enhanced motion vector field, and thermal texture anomaly map to obtain a fused feature vector; perform risk confidence prediction and generation on the fused feature vector to obtain a risk confidence heat map;
[0098] In the embodiment of the present invention, precise spatial and temporal alignment is performed on the temporal spectral stream, enhanced motion vector field, and thermal texture anomaly map, and the data formats are unified. Then, for data of different modalities, pre-trained 3D-CNN and 2D-CNN are respectively used to extract spatio-temporal spectral features, spatial motion features, and thermal texture depth features. Next, a fusion strategy based on the attention mechanism is adopted to learn the importance weights of different modality features, and the weighted features are concatenated and processed through a fully connected layer for feature interaction learning. Finally, through a fully connected layer activated by Sigmoid, the fused feature vector is mapped to a fire risk confidence between 0 and 1, and the confidence value of each pixel is visualized as a risk confidence heat map.
[0099] Step S5: Collect and preprocess the context information of the inspection area according to the sensor network deployment plan to obtain a context information dataset; extract the risk area information from the risk confidence heat map to obtain associated risk area information; perform context-aware warning based on the associated risk area information and the context information dataset to obtain an intelligent warning report;
[0100] In the embodiment of the present invention, environmental data is collected from temperature, humidity, and smoke sensors according to the sensor network deployment plan, and data cleaning is performed to construct a context information dataset. According to the risk confidence heat map generated in step S4, risk areas with a confidence higher than 0.6 are extracted, and their minimum bounding rectangles and center point coordinates are calculated. The risk areas are spatially matched with the context information dataset to associate nearby devices, flammable and explosive items, environmental parameters, and historical alarm records to generate associated risk area information. Different context factors (e.g., flammable and explosive items, electrical equipment, high temperature, low humidity, historical alarms) are weighted according to preset weight values, and a context correction factor is calculated to adjust the original risk confidence to obtain a context-corrected risk confidence. Then, according to the context-corrected risk confidence and combined with the current environmental information, the warning level threshold is dynamically adjusted, and the risk areas are divided into different warning levels such as low, medium, high, and urgent. Finally, for medium, high, and urgent risk areas, visual annotations are made on the visible light images of the temporal spectral stream, and combined with the associated device, environmental parameters, and historical alarm information, the potential causes of fire hazards are analyzed, and an intelligent warning report including the location of the risk area, warning level, potential causes, and disposal suggestions is generated and pushed to relevant personnel by means of text messages and emails.
[0101] Preferably, step S1 includes the following steps:
[0102] Step S11: Deploy multi-spectral cameras in the inspection area and deploy a sensor network to obtain a camera deployment plan and a sensor network deployment plan;
[0103] Step S12: High-speed collect a multi-spectral image sequence according to the camera deployment plan to obtain an original multi-spectral image stream;
[0104] Step S13: Preprocess the original multi-spectral image stream to obtain a corrected multi-spectral image sequence;
[0105] Step S14: Align and fuse the channels of the corrected multi-spectral image sequence to obtain a multi-spectral image stack;
[0106] Step S15: Integrate the temporal spectral information of the multi-spectral image stack to obtain a temporal spectral stream.
[0107] In the embodiment of the present invention, for the inspection area, according to the pre-determined coverage range, potential fire hazard types, and early abnormal spectral features to be captured, a multi-spectral camera including visible light, near-infrared, and short-wave infrared bands and with a resolution of not less than 1920×1080 pixels is selected. Based on the layout map of the inspection area and considering factors such as lighting conditions and potential obstacles, determine the specific installation location, pitch angle, and horizontal angle of the multi-spectral camera to ensure that there is no blind area coverage in the target area. Use a laser rangefinder to accurately measure the distance between the camera installation height and the target area, and use a level to calibrate the camera attitude to ensure the geometric accuracy of image acquisition. Record the camera model, serial number, installation location coordinates, pitch angle, horizontal angle, and initial spectral calibration parameters to form a camera deployment plan. Synchronously carry out sensor network deployment, select wireless sensor nodes including temperature sensors, humidity sensors, and wind speed sensors, and determine the deployment density and location of the sensor nodes according to the environmental characteristics of the inspection area, such as flammable and explosive storage areas and electrical equipment concentration areas. All sensor nodes are connected to the central data acquisition gateway using a star topology and configured with a unique network address and data transmission protocol. Record the model, serial number, deployment location coordinates, and data transmission frequency of each sensor node to form a sensor network deployment plan.
[0108] According to the camera deployment plan formulated in step S11, start the deployed multispectral camera. Set the image acquisition frame rate of the multispectral camera to 10 frames per second to capture early fire signs of rapid change. Configure the image acquisition system to ensure that images in the three spectral channels of visible light, near-infrared, and short-wave infrared can be synchronously acquired, with the time synchronization error controlled within 1 millisecond. Adopt the hardware trigger method or the precise timestamp synchronization mechanism to avoid analysis errors caused by the time difference of image acquisition between channels. Store the acquired raw multispectral images in chronological order in a solid-state drive array in the lossless compressed TIFF format to ensure image quality. Record the timestamp accurate to the millisecond level for each image frame and attach metadata containing camera ID, frame number, and spectral channel information. This process generates a sequence of raw multispectral images that are unprocessed, in chronological order, and contain the three spectral channels of visible light, near-infrared, and short-wave infrared.
[0109] For the raw multispectral image stream generated in step S12, perform preprocessing operations frame by frame. First, for the three spectral channels of visible light, near-infrared, and short-wave infrared, perform radiometric calibration and correction respectively. Using the camera factory calibration parameters, convert the digital gray value of the raw image into a radiance value with physical meaning to eliminate the influence of non-uniform sensor response and environmental light changes. Adopt the dark current correction method to remove the sensor dark current noise, and use an integrating sphere uniform light source to irradiate the calibration image for flat-field correction to eliminate the response difference between sensor pixels. Secondly, adopt a denoising algorithm based on wavelet transform, such as Daubechies wavelet, to suppress the noise of the image in each spectral channel, reduce random noise interference, and improve the signal-to-noise ratio of the image. During the denoising process, the number of wavelet decomposition layers is set to 4 layers, and the soft threshold method is used for threshold processing to retain as much detail information of the image as possible. Finally, if the camera has a slight displacement after deployment, adopt an image registration algorithm based on feature points, such as SIFT or ORB features, to perform geometric correction on the current frame image and the reference frame image, with the correction accuracy controlled at the sub-pixel level. This process generates a sequence of multispectral images that have been radiometrically calibrated, denoised, and geometrically corrected.
[0110] For the corrected multi-spectral image sequence generated in step S13, precise alignment and fusion between spectral channels are performed. First, an image registration algorithm based on maximizing mutual information is adopted. Taking the visible light channel image as the reference, the images of the near-infrared and short-wave infrared channels are precisely registered respectively to ensure pixel-level precise correspondence of the images in the three spectral channels in space. The search window size of the mutual information maximization algorithm is set to 15×15 pixels, and a multi-resolution registration strategy is adopted. Coarse registration is first performed on the low-resolution image, and then fine registration is performed on the high-resolution image. After registration, check and fine-tune the spatial alignment accuracy between the images of different spectral channels to ensure that the position deviation of key feature points is less than 0.5 pixels. Secondly, stack the registered images of the visible light, near-infrared, and short-wave infrared channels in the channel order to form a three-channel image data structure. The image at each time point contains the information of the three spectral channels and can be regarded as a three-dimensional data cube with the data type of float32, which is convenient for subsequent unified processing. This process generates a multi-spectral image stack that is precisely aligned in space.
[0111] For the multi-spectral image stack generated in step S14, integrate it in chronological order to form a four-dimensional data structure including the time dimension. Arrange the multi-spectral image stacks at consecutive time points in ascending order of the acquisition timestamps to construct a four-dimensional tensor, whose dimensions are time, height, width, and channel respectively. The time dimension represents the time sequence of image acquisition, the height and width represent the spatial resolution of the image, and the channel dimension represents different spectral channels (visible light, near-infrared, short-wave infrared). To reduce noise interference in time, a moving average filtering method is used to smooth the time series. The size of the moving window is set to 3 frames, and the radiance values of each pixel point in the respective spectral channels of 3 consecutive frames are averaged to obtain the smoothed spectral information. This process generates the final continuous image sequence including the information of the three spectral channels of visible light, near-infrared, and short-wave infrared, which records the spectral changes of the scene in the time dimension and has the data type of float32.
[0112] Preferably, step S2 includes the following steps:
[0113] Step S21: Perform multi-spectral channel motion estimation on the time-series spectral stream to obtain a channel motion vector map;
[0114] Step S22: Perform motion vector noise suppression and screening on the channel motion vector map to obtain a filtered motion vector map;
[0115] Step S23: Perform spectral characteristic-based motion enhancement on the filtered motion vector map according to the time-series spectral stream to obtain a spectrally enhanced motion vector;
[0116] Step S24: Using the spectral enhancement motion vector as the guiding information, apply the motion magnification algorithm to the temporal spectral flow to obtain the magnified image sequence;
[0117] Step S25: Construct an enhanced motion vector field for the magnified image sequence based on the spectral enhancement motion vector to obtain the enhanced motion vector field.
[0118] As an example of the present invention, refer to Figure 2 As shown, in this example, step S2 includes:
[0119] Step S21: Perform multi-spectral channel motion estimation on the temporal spectral flow to obtain the channel motion vector map;
[0120] In the embodiment of the present invention, for the temporal spectral flow generated in step S15, motion estimation is respectively performed on the image sequences of the visible light, near-infrared, and short-wave infrared spectral channels. A motion estimation algorithm based on pixel-level phase correlation is used to calculate the motion vector of each pixel point between two adjacent frames. Specifically, for the image sequence of each spectral channel, an image block with a size of 32×32 pixels is selected, and a search window is set around the corresponding position of the adjacent frame. The size of the search window is 64×64 pixels. The phase correlation between the current frame image block and all possible position image blocks within the next frame search window is calculated using the fast Fourier transform. The displacement corresponding to the peak of the phase correlation is the motion vector of the central pixel of the image block. The motion vector contains components in the horizontal and vertical directions, and the accuracy reaches the sub-pixel level. For each spectral channel, a vector field with the same size as the image is generated, and each pixel stores the horizontal and vertical direction motion vector information of this pixel between adjacent frames. In this process, a visible light channel motion vector map, a near-infrared channel motion vector map, and a short-wave infrared channel motion vector map are respectively generated.
[0121] Step S22: Perform motion vector noise suppression and screening on the channel motion vector map to obtain the filtered motion vector map;
[0122] In the embodiments of the present invention, for the motion vector maps of the visible light, near-infrared, and short-wave infrared channels generated in step S21, noise suppression and screening are respectively performed. First, the median filtering method is used to suppress the noise in the motion vector maps. For the motion vector map of each channel, a filtering window with a size of 3×3 is selected, and the horizontal and vertical components of all motion vectors within the window are sorted respectively, and the median value is taken as the motion vector component of the current pixel point after filtering. This operation effectively removes isolated and large-amplitude noise motion vectors. Secondly, preliminary screening is performed based on the amplitude of the motion vectors. The amplitude threshold is set to 0.1 pixel. The motion vectors with an amplitude less than this threshold are regarded as tiny motions and are retained; the motion vectors with an amplitude greater than this threshold are regarded as significant motions, such as those caused by the movement of background objects, and are excluded. Finally, for the motion vector map of each channel, the consistency of the motion vectors in the local area is checked. If the direction of the motion vector of a certain pixel point differs from the directions of most of the motion vectors in the surrounding 8-neighborhood by more than 45 degrees, then this motion vector is considered an outlier and is excluded. This process respectively generates the filtered motion vector maps of the visible light channel, near-infrared channel, and short-wave infrared channel.
[0123] Step S23: Perform motion enhancement based on spectral characteristics on the filtered motion vector map according to the temporal spectral flow to obtain the spectrally enhanced motion vectors;
[0124] In the embodiment of the present invention, for the filtered motion vector map generated in step S22, combined with the temporal spectral flow generated in step S15, motion enhancement based on spectral characteristics is performed. First, for each pixel point, the radiance values in the visible light, near-infrared, and short-wave infrared spectral channels are extracted to form a spectral feature vector. A spectral feature model of two early fire characteristics, namely early hot air disturbance and slight smoke, is established. The hot air disturbance model is set to have a high radiance value in the short-wave infrared band, while the slight smoke model is set to have a high scattering intensity in the visible light and near-infrared bands. The cosine similarity between the spectral feature vector of each pixel point and the spectral models of hot air disturbance and slight smoke is calculated. Then, a spectral similarity weight map is constructed. For the motion vector of each pixel point, if the similarity between its corresponding spectral feature and the hot air disturbance model is higher than 0.8, a higher weight, such as 1.5, is assigned to the motion vector; if the similarity with the slight smoke model is higher than 0.8, a weight of 1.2 is assigned; otherwise, a weight of 1.0 is assigned. Next, each motion vector in the filtered motion vector map is multiplied by the corresponding spectral similarity weight value to obtain a spectrally weighted motion vector map. Finally, a temporal filtering method is used to smooth the spectrally weighted motion vector maps of 5 consecutive frames. Exponential smoothing filtering is used, and the smoothing factor is set to 0.2 to enhance persistent small motions and suppress random noise interference. This process generates motion vectors in the visible light channel, near-infrared channel, and short-wave infrared channel after spectral enhancement respectively.
[0125] Step S24: Use the spectrally enhanced motion vectors as guiding information and apply a motion amplification algorithm to the temporal spectral flow to obtain an amplified image sequence;
[0126] In the embodiment of the present invention, for the temporal spectral flow generated in step S15, with the spectrally enhanced motion vectors generated in step S23 as guiding information, a motion amplification algorithm based on optical flow is applied. Specifically, the image sequences in the visible light, near-infrared, and short-wave infrared spectral channels are processed separately. The TV-L1 optical flow algorithm is selected to calculate the dense optical flow field between adjacent frames, and this optical flow field is the spectrally enhanced motion vector. The motion amplification coefficient is set to 10, indicating that the amplitude of small motions is amplified by 10 times. According to the calculated optical flow field, the image warping technique is used to displace the pixels of the current frame according to the indication of the optical flow field to generate the next amplified frame. For example, if the optical flow vector of a certain pixel point indicates that it moves 0.1 pixel to the right, the information of this pixel point and its surrounding pixels is moved to the position 0.1 pixel to the right according to a certain interpolation method. The motion amplification process is performed on the image sequences in the visible light, near-infrared, and short-wave infrared channels respectively. This process generates the visible light image sequence, near-infrared image sequence, and short-wave infrared image sequence after motion amplification.
[0127] Step S25: Construct an enhanced motion vector field for the magnified image sequence based on the spectral enhanced motion vectors to obtain the enhanced motion vector field;
[0128] In the embodiment of the present invention, for the magnified image sequence generated in step S24, an enhanced motion vector field is constructed by combining the spectral enhanced motion vectors generated in step S23. First, the motion vectors of the magnified image sequence are re - estimated. The same motion estimation algorithm based on pixel - level phase correlation as in step S21 is used to calculate the motion vectors of each pixel point between adjacent magnified frames. Since the minute motion has been magnified, the magnitude of the motion vectors estimated at this time is larger and easier to detect. Then, the spectral enhanced motion vectors are fused with the re - estimated motion vectors. For each pixel point, if the magnitude of its spectral enhanced motion vector is greater than the magnitude of the re - estimated motion vector, the spectral enhanced motion vector is selected as the final motion vector; otherwise, the re - estimated motion vector is selected. This fusion strategy aims to use spectral information to guide more accurate motion estimation. Finally, the final motion vector information is organized into an enhanced motion vector field. For each pixel point, the horizontal and vertical components of its motion vector are stored. This process generates an enhanced motion vector field including three channels of visible light, near - infrared, and short - wave infrared.
[0129] Preferably, step S23 includes the following steps:
[0130] Step S231: Calculate the spectral feature response maps for the temporal spectral flow to obtain a set of spectral feature response maps;
[0131] Step S232: Construct a spectral similarity weight map based on the set of spectral feature response maps to obtain a set of spectral similarity weight maps;
[0132] Step S233: Fuse the spectral weights and motion vectors of the filtered motion vector map according to the set of spectral similarity weight maps to obtain a spectrally weighted motion vector map;
[0133] Step S234: Smooth and accumulate the motion vectors in the time domain for the spectrally weighted motion vector map to obtain a temporally smoothed motion vector map;
[0134] Step S235: Enhance the motion vectors according to the temporally smoothed motion vector map and the filtered motion vector map to obtain the spectral enhanced motion vectors.
[0135] In the embodiments of the present invention, for the temporal spectral stream generated in step S15, the spectral responses with predefined fire characteristics are calculated pixel by pixel. First, the radiance values of each pixel in the three channels of visible light, near-infrared, and short-wave infrared are extracted to form a three-dimensional spectral feature vector. Standard spectral response models for three fire characteristics, namely early smoldering, open fire, and overheating, are established. The early smoldering model is set such that the reflectance in the near-infrared band is higher than that in the visible light and short-wave infrared bands; the open fire model is set such that the radiation intensity in the visible light band is the highest, followed by the near-infrared band, and the short-wave infrared intensity is lower; the overheating model is set such that the radiation intensity in the short-wave infrared band is significantly higher than that in the visible light and near-infrared bands. Then, the cosine similarity between the spectral feature vector of each pixel and the three standard spectral response models is calculated. The range of the cosine similarity value is from -1 to 1, and the closer the value is to 1, the higher the similarity. For each pixel, the similarity values with the early smoldering, open fire, and overheating models are calculated respectively. Finally, the similarity values of each pixel with the three fire characteristics are stored respectively to generate three spectral feature response maps, corresponding to the early smoldering response map, the open fire response map, and the overheating response map respectively. The pixel value range of each response map is from -1 to 1.
[0136] For the spectral feature response map set generated in step S231, that is, the early smoldering response map, the open fire response map, and the overheating response map, a spectral similarity weight map is constructed. First, three weight mapping functions are set, corresponding to the three fire characteristics of early smoldering, open fire, and overheating respectively. For early smoldering, when the pixel value of its response map is greater than 0.6, the weight mapping function outputs 1.2; when the pixel value is less than 0.3, it outputs 0.8; when it is between 0.3 and 0.6, linear interpolation is used for output. For open fire, when the pixel value of its response image is greater than 0.8, the weight mapping function outputs 1.5; when it is less than 0.2, it outputs 0.7; linear interpolation is used for intermediate values. For overheating, when the pixel value of its response image is greater than 0.7, the weight mapping function outputs 1.3; when it is less than 0.4, it outputs 0.9; linear interpolation is used for intermediate values. Then, the pixel values of each spectral feature response map are respectively substituted into the corresponding weight mapping function to calculate the spectral similarity weight value of each pixel. Finally, the weight values of each pixel for early smoldering, open fire, and overheating are stored respectively to generate three spectral similarity weight maps, corresponding to the early smoldering weight map, the open fire weight map, and the overheating weight map respectively. The pixel value range of each weight map is set according to the mapping function.
[0137] For the filtered motion vector map generated in step S22 and the spectral similarity weight map set generated in step S232, perform the fusion of spectral weights and motion vectors. First, for each motion vector (including horizontal and vertical components) in the filtered motion vector map, read the weight values of the corresponding pixels from the early smoldering weight map, the flaming weight map, and the overheating weight map respectively. Then, adopt a weighted average fusion strategy. Multiply the filtered motion vector by the three weight values respectively to obtain three weighted motion vectors. The final spectrally weighted motion vector is obtained by taking the weighted average of these three weighted motion vectors. For example, let the early smoldering weight be w1, the flaming weight be w2, the overheating weight be w3, and the filtered motion vector be V, then the spectrally weighted motion vector V' = (w1 * V + w2 * V + w3 * V) / (w1 + w2 + w3). If all three weight values are zero, the spectrally weighted motion vector remains unchanged as the filtered motion vector. Finally, store the fused motion vector of each pixel to generate a spectrally weighted motion vector map. This vector map contains two components, horizontal and vertical.
[0138] For the sequence of spectrally weighted motion vector maps generated in step S233, perform temporal motion vector smoothing and accumulation. Use the Kalman filter algorithm to smooth the motion vectors of each pixel over time. First, establish a uniform motion model, assuming that the motion state (position and velocity) of the pixel changes linearly over time. The state vector includes the x coordinate, y coordinate, x-direction velocity, and y-direction velocity of the pixel. The observation vector includes the spectrally weighted motion vector of the pixel in the current frame. Then, initialize the state estimate and covariance matrix of the Kalman filter. For each frame of the image, predict the state of the current frame based on the state estimate of the previous frame and the motion model. Next, update the state estimate according to the spectrally weighted motion vector of the current frame. The process noise covariance matrix and observation noise covariance matrix of the Kalman filter are set according to experience. Through the Kalman filter, the random noise in the motion vectors can be effectively smoothed out, and the persistent motion information can be accumulated. Finally, store the motion vectors after filtering for each frame to generate a sequence of temporally smoothed motion vector maps. This vector map contains two components, horizontal and vertical.
[0139] For the time-domain smoothed motion vector map generated in step S234 and the filtered motion vector map generated in step S22, perform final motion vector enhancement. First, calculate the magnitudes of the time-domain smoothed motion vector map and the filtered motion vector map. For each pixel, calculate the magnitude values in its time-domain smoothed motion vector and the filtered motion vector respectively. Then, set an enhancement threshold, for example, 0.05 pixels. If the magnitude of the time-domain smoothed motion vector of a certain pixel is greater than this threshold and its direction is the same as that of the filtered motion vector (the direction angle is less than 45 degrees), then enhance the filtered motion vector. The enhancement method is to magnify the magnitude of the filtered motion vector by 1.2 times while keeping the direction unchanged. If the magnitude of the time-domain smoothed motion vector is less than the threshold, or the direction is inconsistent with that of the filtered motion vector, then the spectral enhanced motion vector remains unchanged as the filtered motion vector. Finally, store the enhanced motion vector of each pixel to generate a spectral enhanced motion vector map. This vector map contains two components: horizontal and vertical.
[0140] Preferably, step S3 includes the following steps:
[0141] Step S31: Extract an infrared image sequence from the temporal spectral stream and perform inverse geometric transformation using the enhanced motion vector field to obtain an aligned infrared image sequence;
[0142] Step S32: Perform preprocessing on the aligned infrared image sequence to obtain a preprocessed thermal image sequence;
[0143] Step S33: Extract thermal texture features from the preprocessed thermal image sequence to obtain a thermal texture feature map;
[0144] Step S34: Perform analysis on the dynamic changes of the thermal texture of the thermal texture feature map to obtain a thermal texture change map;
[0145] Step S35: Locate and characterize thermal texture anomalies in the preprocessed thermal image sequence according to the thermal texture change map to obtain a thermal texture anomaly map.
[0146] As an example of the present invention, refer to Figure 3 As shown, in this example, step S3 includes:
[0147] Step S31: Extract an infrared image sequence from the temporal spectral stream and perform inverse geometric transformation using the enhanced motion vector field to obtain an aligned infrared image sequence;
[0148] In the embodiment of the present invention, for the temporal spectral stream generated in step S15, an image sequence of the short-wave infrared channel is extracted as the initial infrared image sequence. The short-wave infrared band is sensitive to the thermal radiation generated by early fires. Then, for the enhanced motion vector field generated in step S25, an inverse geometric transformation is performed on the initial infrared image sequence. Specifically, for each frame in the initial infrared image sequence, according to the motion vector of the corresponding pixel in the enhanced motion vector field, the pixel position is inversely offset. For example, if the enhanced motion vector field indicates that a certain pixel has moved 2 pixels to the right and 1 pixel down from the previous frame to the current frame, then in the inverse geometric transformation, the pixel value at the pixel position in the current frame is assigned to the position that is 2 pixels to the left and 1 pixel up in the previous frame. The inverse geometric transformation uses the bilinear interpolation method for pixel value filling to reduce image distortion. This operation aims to eliminate the image deformation introduced by the previous motion magnification process and realign the infrared image sequence with the original scene spatially. This process generates an infrared image sequence that is spatially aligned with the original scene.
[0149] Step S32: Perform thermal image preprocessing on the aligned infrared image sequence to obtain a preprocessed thermal image sequence;
[0150] In the embodiment of the present invention, for the aligned infrared image sequence generated in step S31, a thermal image preprocessing operation is performed. First, the Gaussian filtering method is used to suppress the noise in the infrared image. A Gaussian filtering kernel with a size of 5×5 and a standard deviation set to 1.5 is selected to smooth each frame of the infrared image and reduce random noise interference. Second, the histogram equalization method is used to enhance the contrast of the thermal image. For each frame of the infrared image, the histogram of its pixel gray values is calculated, and the cumulative distribution function is calculated. Then, according to the cumulative distribution function, the pixel gray values are remapped so that the gray distribution of the image is more uniform, thereby enhancing the contrast of the image and making the temperature difference more obvious. Finally, the pixel gray values of the infrared image are linearly mapped to the range of 0 to 255 for normalization processing to facilitate subsequent texture feature calculation. This process generates a preprocessed thermal image sequence that has been subjected to noise suppression, contrast enhancement, and normalization processing.
[0151] Step S33: Extract thermal texture features from the preprocessed thermal image sequence to obtain a thermal texture feature map;
[0152] In the embodiment of the present invention, for the preprocessed thermal image sequence generated in step S32, thermal texture features are extracted. The local binary pattern (LBP) operator is used to extract the local texture information of the image. Specifically, for each frame of the thermal image, with each pixel as the center, a circular neighborhood with a radius of 3 pixels is selected, and 8 sampling points are evenly selected within the neighborhood. The gray value of the neighborhood sampling point is compared with the gray value of the central pixel. If the gray value of the neighborhood sampling point is greater than or equal to the gray value of the central pixel, it is marked as 1, otherwise it is marked as 0. The comparison results of the 8 sampling points are arranged in sequence to form an 8-bit binary number, which is the LBP value of the central pixel. Calculate the LBP value of each pixel in the entire image to generate an LBP texture image. The LBP value ranges from 0 to 255, and different LBP values correspond to different local texture patterns. This process generates a sequence of LBP thermal texture feature maps.
[0153] Step S34: Perform dynamic thermal texture change analysis on the thermal texture feature map to obtain a thermal texture change map;
[0154] In the embodiment of the present invention, for the sequence of LBP thermal texture feature maps generated in step S33, dynamic thermal texture change analysis is performed. First, calculate the LBP histograms of two adjacent frames of LBP thermal texture feature maps. For each frame of LBP image, count the frequencies of the LBP values from 0 to 255 to obtain a 256-dimensional LBP histogram. Then, calculate the Bhattacharyya distance between the two adjacent LBP histograms. The Bhattacharyya distance is used to measure the similarity between two probability distributions, and its value range is from 0 to 1. The smaller the value, the more similar the two histograms are and the smaller the texture change. The calculation formula is: Bhattacharyya distance = 1 - sqrt(the sum of H1[i] * H2[i]), where H1 and H2 respectively represent the LBP histograms of two adjacent frames. Finally, use the calculated Bhattacharyya distance value as the thermal texture change value of the current frame to generate a thermal texture change map. The pixel value range of the thermal texture change map is from 0 to 1, and the larger the value, the more significant the texture change. This process generates a sequence of thermal texture change maps.
[0155] Step S35: Perform thermal texture anomaly localization and characterization on the preprocessed thermal image sequence according to the thermal texture change map to obtain a thermal texture anomaly map;
[0156] In the embodiments of the present invention, for the sequence of thermal texture change maps generated in step S34 and the sequence of preprocessed thermal images generated in step S32, thermal texture anomaly localization and characterization are performed. First, a thermal texture change threshold is set, for example, 0.2. Pixel points in the thermal texture change map with pixel values greater than this threshold are marked as anomaly points to generate a binary anomaly mask. Then, connected component analysis is performed on the binary anomaly mask, and adjacent anomaly pixel points are divided into the same anomaly region, and a unique label is assigned to each anomaly region. Next, the pixel set corresponding to each anomaly region is extracted from the sequence of preprocessed thermal images, and the average temperature, maximum temperature, minimum temperature, and temperature standard deviation of each anomaly region are calculated. At the same time, the LBP histogram of each anomaly region is calculated, and the mean, variance, skewness, and kurtosis of the LBP histogram are extracted as texture features. Finally, the anomaly regions are classified according to predefined rules. For example, if the average temperature of an anomaly region is more than 5 degrees Celsius higher than the ambient temperature and the variance of the LBP histogram is large, it is marked as a potential overheating point; if the texture change value of an anomaly region is continuously high but the temperature increase is not obvious, it is marked as a potential hot air disturbance. This process generates a thermal texture anomaly map including the positions, temperature features, and texture features of the anomaly regions, and marks different types of anomalies.
[0157] Preferably, step S34 includes the following steps:
[0158] Step S341: Calculate a texture difference feature map for the thermal texture feature map to obtain a texture difference feature map;
[0159] Step S342: Establish a motion amplitude feature map for the enhanced motion vector field to obtain a motion amplitude feature map;
[0160] Step S343: Perform texture difference and motion amplitude fusion on the texture difference feature map and the motion amplitude feature map to obtain a fusion feature score map;
[0161] Step S344: Perform time series anomaly detection on the fusion feature score map to obtain an anomaly score map;
[0162] Step S345: Generate a thermal texture change map according to the anomaly score map to obtain a thermal texture change map.
[0163] In the embodiments of the present invention, for the sequence of LBP thermal texture feature maps generated in step S33, the difference in texture features between adjacent frames is calculated. The specific operation is as follows: for two adjacent frames of images in the LBP thermal texture feature map sequence, the absolute difference in the LBP values of the corresponding pixel points is calculated. For example, let the LBP value of the pixel point (x, y) in the previous frame image be LBP1(x, y), and the LBP value of the pixel point (x, y) in the next frame image be LBP2(x, y), then the texture difference value of this pixel point is |LBP1(x, y) - LBP2(x, y)|. Calculate the texture difference values of all pixel points in the entire image to generate a texture difference feature map. The pixel value range of the texture difference feature map is from 0 to 255, and the larger the value, the more significant the texture change near the pixel point. This process generates a sequence of texture difference feature maps.
[0164] For the enhanced motion vector field generated in step S25, the motion amplitude of each pixel point is calculated. The specific operation is as follows: for each pixel point in the enhanced motion vector field, extract its corresponding motion vector, and this motion vector includes two components: horizontal and vertical. Use the Pythagorean theorem to calculate the amplitude value of this motion vector, that is, motion amplitude = sqrt(horizontal component^2 + vertical component^2). Calculate the motion amplitude values of all pixel points in the entire image to generate a motion amplitude feature map. The pixel value range of the motion amplitude feature map is from 0 to the length of the image diagonal, and the larger the value, the larger the motion amplitude of the pixel point. This process generates a sequence of motion amplitude feature maps.
[0165] For the sequence of texture difference feature maps generated in step S341 and the sequence of motion amplitude feature maps generated in step S342, the fusion of texture difference and motion amplitude is performed. The specific operation is as follows: for the texture difference feature map and the motion amplitude feature map at the same moment, the texture difference value and the motion amplitude value of the corresponding pixel points are weighted and fused. Set the weight of the texture difference to 0.6 and the weight of the motion amplitude to 0.4. The fused feature score is: fused feature score = 0.6 * texture difference value + 0.4 * motion amplitude value. Calculate the fused feature scores of all pixel points in the entire image to generate a fused feature score map. The pixel value range of the fused feature score map is determined according to the ranges of the texture difference value and the motion amplitude value, and the larger the value, the more significant the texture change and the larger the motion amplitude near the pixel point. This process generates a sequence of fused feature score maps.
[0166] Perform time series anomaly detection on the sequence of fused feature score maps generated in step S343. The specific operation is as follows: for each pixel point in the sequence of fused feature score maps, extract its fused feature score values in the past 5 frames to form a time series. An anomaly detection method based on a sliding window is adopted. Set the sliding window size to 5 frames, and calculate the mean and standard deviation of the fused feature score values of the current frame pixel point in the past 5 frames. Then, calculate the difference between the fused feature score value of the current frame and the mean score of the past 5 frames, and divide this difference by the standard deviation of the scores of the past 5 frames to obtain a Z-score value. The larger the Z-score value, the farther the fused feature score of the current frame deviates from the average level in the past period of time, and the more likely it is to be an anomaly. Set the anomaly threshold to 2.0. If the Z-score value of the current frame is greater than 2.0, it is considered that there is an anomaly at this pixel point. Store the Z-score values of all pixel points in the entire image to generate an anomaly score map. The pixel value of the anomaly score map represents the possibility of thermal texture dynamic anomaly at this pixel point. This process generates a sequence of anomaly score maps.
[0167] Generate the final thermal texture change map for the sequence of anomaly score maps generated in step S344. The specific operation is as follows: directly output the anomaly score map of the current frame as the thermal texture change map. The pixel value of the thermal texture change map is the anomaly score of this pixel point. The higher the score, the more significant the dynamic change of the thermal texture in this area, and the more likely it is to be a sign of an early fire. The anomaly score map can be normalized, for example, map the score value to the gray scale range of 0 to 255 for convenient subsequent visualization and analysis. This process generates the final thermal texture change map.
[0168] Preferably, step S35 includes the following steps:
[0169] Step S351: Generate a binary anomaly mask for the thermal texture change map to obtain a binary anomaly mask;
[0170] Step S352: Perform component labeling connection on the binary anomaly mask to obtain a connected component label map;
[0171] Step S353: Extract the anomaly regions from the preprocessed thermal image sequence according to the connected component label map to obtain a set of anomaly region image patches;
[0172] Step S354: Calculate the regional temperature statistical features for the set of anomaly region image patches to obtain a regional temperature feature table;
[0173] Step S355: Calculate the regional texture features for the set of anomaly region image patches to obtain a regional texture feature table;
[0174] Step S356: Classify and label the anomaly patterns according to the regional temperature feature table and the regional texture feature table to obtain a labeled anomaly region map;
[0175] Step S357: Generate a thermal texture anomaly map based on the marked anomaly region map, the regional temperature feature table, and the regional texture feature table to obtain the thermal texture anomaly map.
[0176] In the embodiment of the present invention, for the thermal texture change map generated in step S345, a binary anomaly mask is generated. The specific operation is as follows: Set an anomaly threshold, which is set to 0.7 according to historical data statistical analysis or expert experience. Traverse each pixel point in the thermal texture change map. If the thermal texture change value of the pixel point is greater than or equal to 0.7, set the value at the corresponding position in the binary anomaly mask to 255; otherwise, set it to 0. This operation marks the regions with significant changes in the thermal texture change map. The binary anomaly mask is a single-channel image, and the pixel values are only 0 or 255. This process generates the binary anomaly mask.
[0177] For the binary anomaly mask generated in step S351, connected component labeling is performed. The specific operation is as follows: Adopt an eight-neighborhood connection method and traverse each pixel point in the binary anomaly mask. If the current pixel value is 255 and there are pixel points with a pixel value of 255 in its eight neighborhoods, mark them as the same connected component and assign the same label value. For isolated anomaly pixel points that are not adjacent to any other anomaly pixel points, also assign a unique label value. The label values start from 1 and increase. The connected component labeling map is a single-channel image, and different pixel values represent different anomaly regions. This process generates the connected component labeling map.
[0178] For the connected component labeling map generated in step S352 and the preprocessed thermal image sequence generated in step S32, the extraction of anomaly regions is performed. The specific operation is as follows: Traverse each label value (representing an anomaly region) in the connected component labeling map. For each label value, in the current frame of the preprocessed thermal image sequence, extract all pixel points whose pixel values are equal to the label value. These pixel points form an anomaly region. Crop out the image block corresponding to the pixel points of each anomaly region in the preprocessed thermal image to form an independent image block. The size and shape of each image block depend on the size and shape of the anomaly region. This process generates a set containing multiple anomaly region image blocks.
[0179] For the set of abnormal region image patches generated in step S353, calculate the temperature statistical features of each abnormal region. The specific operation is to traverse each abnormal region image patch and calculate the average value, maximum value, minimum value, and standard deviation of the temperature values of all pixels within the image patch. The temperature value is obtained by converting the pixel value of the infrared image, and the conversion formula is determined according to the camera calibration parameters. For example, the average temperature calculation method is to add up the temperature values of all pixels in the region and then divide by the total number of pixels. The maximum and minimum values directly take the maximum and minimum temperature values of the pixels in the region. The standard deviation reflects the degree of dispersion of the temperature distribution within the region. Record the average temperature, maximum temperature, minimum temperature, and standard deviation of each abnormal region in a table, where each row of the table corresponds to an abnormal region and each column corresponds to a temperature statistical feature. This process generates a regional temperature feature table.
[0180] For the set of abnormal region image patches generated in step S353, calculate the texture features of each abnormal region. The specific operation is to traverse each abnormal region image patch and calculate the local binary pattern (LBP) histogram of the image patch. The calculation method of the LBP histogram is similar to that in step S341, and the frequency of each LBP value appearing within the image patch is counted. Then, extract the statistical features of the LBP histogram, including the mean, variance, skewness, and kurtosis. The mean reflects the central position of the LBP value distribution, the variance reflects the degree of dispersion of the LBP value distribution, the skewness reflects the symmetry of the LBP value distribution, and the kurtosis reflects the sharpness of the LBP value distribution. Record the statistical features of the LBP histogram of each abnormal region in a table, where each row of the table corresponds to an abnormal region and each column corresponds to a texture statistical feature. This process generates a regional texture feature table.
[0181] For the regional temperature feature table generated in step S354 and the regional texture feature table generated in step S355, classify and label the abnormal patterns. The specific operation is to, for each abnormal region, combine its corresponding temperature feature and texture feature into a feature vector. Use a pre-trained classification model (such as a support vector machine or a random forest) to classify the feature vector and determine which predefined abnormal pattern the abnormality belongs to, such as overheating, early smoke, or flame. The training data of the classification model comes from historical fire cases or simulated fire experiment data. According to the classification result, assign a corresponding label to each abnormal region, such as "overheating", "smoke", or "flame". Then, in the connected component labeling map, label the pixel points belonging to the same abnormal region with the corresponding color or assign the corresponding label value. This process generates a labeled abnormal region map.
[0182] Generate a final thermal texture anomaly map based on the marked anomaly region map generated in step S356, the region temperature feature table generated in step S354, and the region texture feature table generated in step S355. The specific operation is to overlay the marked anomaly regions on the preprocessed thermal image and use different colors or bounding boxes to indicate different types of anomaly regions. For example, overheated regions are indicated in red, smoke regions are indicated in blue, and flame regions are indicated in yellow. Add a legend on the side or below the map to explain the anomaly types represented by different colors or markings. At the same time, display the detailed information of each anomaly region in the map. For example, when the mouse hovers over an anomaly region, information such as the average temperature, maximum temperature, and mean value of the LBP histogram of that region pops up. The temperature features and texture features of the anomaly regions can be added to the map in the form of a table. This process generates a thermal texture anomaly map containing the positions, types, and detailed feature information of the anomaly regions.
[0183] Preferably, step S4 includes the following steps:
[0184] Step S41: Align the multi-modal feature data of the temporal spectral flow, enhanced motion vector field, and thermal texture anomaly map to obtain an aligned multi-modal feature set;
[0185] Step S42: Extract multi-modal deep features from the aligned multi-modal feature set to obtain multi-modal deep features;
[0186] Step S43: Perform feature-level fusion on the multi-modal deep features and conduct feature interaction learning to obtain a fused feature vector;
[0187] Step S44: Predict and generate the risk confidence level for the fused feature vector to obtain a risk confidence level thermal map.
[0188] In the embodiments of the present invention, alignment operations for multi-modal feature data are performed on the temporal spectral stream generated in step S15, the enhanced motion vector field generated in step S25, and the thermal texture anomaly map generated in step S357. First, based on the spatial resolution and pixel coordinate system of the visible light channel image of the temporal spectral stream, the spatial resolutions of the enhanced motion vector field and the thermal texture anomaly map are unified to this benchmark. For the enhanced motion vector field, the bilinear interpolation method is used to adjust the dimension of the motion vector field to be consistent with the visible light image. For the thermal texture anomaly map, if its resolution is different from that of the visible light image, the nearest neighbor interpolation method is used to adjust its resolution. Secondly, ensure that the data of different modalities are synchronized in time. Assume that the frame rate of the temporal spectral stream is 10 frames per second, and the enhanced motion vector field and the thermal texture anomaly map are also generated at the same frame rate, then no additional time synchronization operation is required. If the frame rates of the data of different modalities are different, time interpolation or decimation operations are required to align them in time. Finally, the aligned multi-modal data are combined into a multi-channel data set. For the image at each time point, it includes the image data of three spectral channels, namely visible light, near-infrared, and short-wave infrared, as well as the corresponding enhanced motion vector field (including two components, horizontal and vertical) and the thermal texture anomaly map (including anomaly class labels). This process generates the aligned multi-modal feature set.
[0189] For the aligned multi-modal feature set generated in step S41, depth features of different modalities are respectively extracted. First, for the temporal spectral stream, a pre-trained three-dimensional convolutional neural network (3D-CNN), such as ResNet3D-18, is used to extract spatio-spectral features. This network is pre-trained on the ImageNet video data set and then fine-tuned on this inspection data set. The multi-spectral image at each time point is input into the 3D-CNN, and the output of the last convolutional layer is extracted as the spectral depth feature. Secondly, for the enhanced motion vector field, a two-dimensional convolutional neural network (2D-CNN), such as VGG16, is used to extract spatial motion features. The horizontal and vertical components of the enhanced motion vector field are used as two input channels and input into the 2D-CNN, and the output of the last pooling layer is extracted as the motion depth feature. This network is also pre-trained and fine-tuned on the ImageNet data set. Finally, for the thermal texture anomaly map, another 2D-CNN, such as MobileNetV2, is used to extract thermal texture depth features. The thermal texture anomaly map is used as the input, and the output of the last convolutional layer is extracted as the thermal texture depth feature. This process generates spectral depth features, motion depth features, and thermal texture depth features.
[0190] For the spectral depth features, motion depth features, and thermal texture depth features generated in step S42, perform feature-level fusion and interactive learning. Adopt a fusion strategy based on the attention mechanism. First, input the depth features of the three modalities into three independent attention modules respectively. Each attention module contains a fully connected layer and a Softmax activation function to learn the importance weights of the features of this modality. Then, multiply the depth features of each modality by their corresponding attention weights to obtain the weighted feature representations. Next, concatenate the weighted depth features of the three modalities to obtain a fused feature vector. To learn the interaction information between different modality features, add several fully connected layers after the concatenated feature vector and use the Dropout regularization method to prevent overfitting. Through end-to-end training, let the network learn the correlation between different modality features and assign different weights to different modality features to highlight the features that are more important for fire hazard judgment. This process generates a fused feature vector.
[0191] For the fused feature vector generated in step S43, predict the fire risk confidence. The specific operation is to connect a fully connected layer at the end of the fused feature vector. This fully connected layer contains an output unit and uses the Sigmoid activation function to map the fused feature vector to a scalar value between 0 and 1, and this value is the confidence of the fire risk. The closer the confidence value is to 1, the higher the possibility that there is a fire hazard in this area. Then, map the risk confidence value of each pixel point to a heat map. The heat map uses pseudo-color coding. For example, map the confidence value 0 to blue, the confidence value 1 to red, and perform linear interpolation for intermediate values. Fill the risk confidence value of each pixel point into the corresponding position of the heat map to generate a risk confidence heat map. This process generates a risk confidence heat map.
[0192] Preferably, step S5 includes the following steps:
[0193] Step S51: Collect and preprocess the situational information of the inspection area according to the sensor network deployment plan to obtain a situational information dataset;
[0194] Step S52: Extract the risk areas from the risk confidence heat map to obtain risk area data; match the spatial information of the risk area data and the situational information dataset, and perform risk area association to obtain associated risk area information;
[0195] Step S53: Perform situational factor weighting according to the associated risk area information to obtain situational factor weighted data; perform risk correction according to the situational factor weighted data to obtain a situation-corrected risk confidence;
[0196] Step S54: performing warning level classification according to the scenario-corrected risk confidence level to obtain warning level classification information; dynamically adjusting the threshold of the warning level classification information according to the scenario information data set to obtain warning level information;
[0197] Step S55: Generate and push an intelligent warning report according to the warning level information, the associated risk area information and the time series spectrum flow to obtain an intelligent warning report.
[0198] In the embodiment of the present invention, according to the sensor network deployment scheme formulated in step S11, environmental parameter data is collected in real time from the temperature sensor, humidity sensor and smoke sensor deployed in the inspection area. The data collection frequency is once per minute. The collected raw data is preprocessed, firstly, data cleaning is performed to remove abnormal data that exceeds the sensor range or is obviously wrong, such as data with a temperature value exceeding 100 degrees Celsius or a negative humidity value. Secondly, data verification is performed, and the data collected three times in a row are smoothed using a sliding average filtering method, and the window size is 3. Then the data is formatted, and the collected temperature, humidity and smoke concentration values are converted to floating point types, and corresponding timestamp information is added. At the same time, the pre-stored static scene semantic information of the inspection area is loaded, including the CAD Vector Drawings drawings of the inspection area, the equipment layout information (including equipment name, model, location coordinates) and the distribution map of flammable and explosive items (including item name, quantity, storage location). The CAD drawings are converted into regional information represented by polygons, and spatially associated with the equipment layout and the distribution information of flammable and explosive items to determine the area where each device and flammable and explosive item is located. In addition, the historical fire alarm records of the inspection area in the past year are retrieved from the historical database, including the alarm time, alarm type and alarm location. This process generates a context information dataset containing real-time environmental parameters, static scene semantic information and historical alarm records.
[0199] For the risk confidence heat map generated in step S44, extract the risk areas. Set the risk threshold to 0.6, and mark the pixel points with pixel values greater than or equal to 0.6 in the risk confidence heat map as high-risk pixel points. Then, use the eight-neighborhood connected component analysis algorithm to connect adjacent high-risk pixel points into a risk area. Calculate the minimum bounding rectangle for each independent risk area, and record its center point coordinates, width, and height. This process generates risk area data containing multiple risk areas. Next, perform spatial information matching between each extracted risk area and the situation information dataset generated in step S51. Using the center point coordinates of the risk area, determine whether the risk area is located within a predefined area (e.g., flammable and explosive storage area, electrical equipment area). If the spatial distance between the risk area and a device or flammable and explosive item is less than 1 meter, it is considered that the risk area is associated with that device or flammable and explosive item. At the same time, obtain the real-time temperature, humidity, and smoke concentration values corresponding to the risk area, as well as the historical alarm records of this area. This process generates associated risk area information, including the location, risk confidence, and related environmental parameters, scenario semantic information, and historical alarm records of each risk area.
[0200] For the associated risk area information generated in step S52, perform weighting of the situation factors. Set weight values for different situation factors, and the weight value range is from 0 to 1, indicating the degree of influence of this factor on the fire risk. For example, if the risk area is located in the flammable and explosive storage area, set the weight of flammable and explosive items in this area to 0.8; if it is located in the electrical equipment area, set the weight of electrical equipment to 0.7; if the current environmental temperature is higher than 35 degrees Celsius, set the weight of high temperature to 0.6; if the humidity is lower than 30%, set the weight of low humidity to 0.5; if there has been a fire alarm in this area in the past year, set the weight of historical alarms to 0.9. This process generates situation factor weighting data containing each risk area and its corresponding situation factor weights. Then, correct the original risk confidence according to the situation factor weighting data. For each risk area, calculate its situation correction factor, and the calculation formula is: situation correction factor = 1 + (weight of flammable and explosive items + weight of electrical equipment + weight of high temperature + weight of low humidity + weight of historical alarms) / 5. Multiply the original risk confidence by the situation correction factor to obtain the situation-corrected risk confidence. The situation-corrected risk confidence value may be greater than 1. This process generates situation-corrected risk confidence.
[0201] For the situation correction risk confidence generated in step S53, the early warning levels are divided. Four early warning levels are set: low, medium, high, and urgent. The division thresholds for the early warning levels are as follows: If the situation correction risk confidence is less than 0.7, it is a low risk; if it is greater than or equal to 0.7 and less than 0.85, it is a medium risk; if it is greater than or equal to 0.85 and less than 1.0, it is a high risk; if it is greater than or equal to 1.0, it is an urgent risk. This process generates early warning level division information including each risk area and its corresponding initial early warning level. Then, according to the situation information dataset generated in step S51, the thresholds for early warning level division are dynamically adjusted. If the average temperature in the current inspection area is higher than 30 degrees Celsius, the thresholds for all early warning levels are reduced by 0.05; if the wind speed in the current area is greater than 5 m / s, the thresholds for high risk and urgent risk are reduced by 0.1. The dynamically adjusted thresholds are used for the final early warning level judgment. This process generates early warning level information including each risk area and its final early warning level.
[0202] Generate an intelligent early warning report based on the early warning level information generated in step S54, the associated risk area information generated in step S52, and the time-series spectral flow generated in step S15. First, according to the early warning level information, determine the risk areas for which early warning reports need to be generated. Only generate early warning reports for risk areas with early warning levels of medium, high, or urgent. Then, for each risk area for which an early warning report needs to be generated, use a red border to Highlight the location of the risk area on the visible light image of the time-series spectral flow and label its early warning level. Analyze the possible causes of fire hazards in the risk area based on the equipment or flammable and explosive material information related to the risk area recorded in the associated risk area information, as well as the current environmental parameters and historical alarm records, such as "equipment overheating risk", "flammable material leakage risk", etc. Give corresponding disposal suggestions according to the risk level and potential causes, such as "recommend strengthening monitoring", "recommend sending someone to check immediately", "recommend activating the emergency plan". Integrate information such as the location of the risk area, early warning level, potential cause analysis, and disposal suggestions into a structured report, and the report format is JSON or XML. Finally, push the generated intelligent early warning report to the preset fire safety management personnel and inspection personnel via text message and email. This process generates an intelligent early warning report.
[0203] Preferably, the present invention also provides a fire safety inspection system based on machine vision recognition for performing the fire safety inspection method based on machine vision recognition as described above. The fire safety inspection system based on machine vision recognition includes:
[0204] A multi-spectral dynamic perception module for deploying inspection equipment in the inspection area to obtain a camera deployment plan and a sensor network deployment plan; performing multi-spectral dynamic perception according to the camera deployment plan to obtain a time-series spectral flow;
[0205] A micro-motion amplification module, which is used to perform multi-spectral channel motion estimation on the time-series spectral stream, suppress the motion vector noise, and obtain a filtered motion vector map; perform motion enhancement based on spectral characteristics on the filtered motion vector map to obtain a spectrally enhanced motion vector; apply a motion amplification algorithm to the time-series spectral stream according to the spectrally enhanced motion vector, and construct an enhanced motion vector field to obtain an enhanced motion vector field;
[0206] A thermal texture anomaly analysis module, which is used to extract thermal texture features from the time-series spectral stream to obtain a thermal texture feature map; perform dynamic change analysis of the thermal texture feature map to obtain a thermal texture change map; perform thermal texture anomaly localization and characterization on the preprocessed thermal image sequence according to the thermal texture change map to obtain a thermal texture anomaly map;
[0207] A multi-modal feature fusion module, which is used to perform multi-modal feature fusion on the time-series spectral stream, the enhanced motion vector field, and the thermal texture anomaly map to obtain a fused feature vector; perform risk confidence prediction and generation on the fused feature vector to obtain a risk confidence heat map;
[0208] A context-aware warning module, which is used to collect and preprocess context information of the inspection area according to the sensor network deployment scheme to obtain a context information data set; extract risk area information from the risk confidence heat map to obtain associated risk area information; perform context-aware warning according to the associated risk area information and the context information data set to obtain an intelligent warning report.
[0209] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the application document are intended to be included in the present invention.
[0210] The above are only specific embodiments of the present invention, which enable those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A fire safety inspection method based on machine vision recognition, characterized in that: The following steps are involved: Step S1: Deploy inspection equipment in the inspection area to obtain a camera deployment plan and a sensor network deployment plan; perform multi-spectral dynamic perception according to the camera deployment plan to obtain a time-series spectral stream; Step S2: performing multi-spectral channel motion estimation on the time series spectral stream and performing motion vector noise suppression to obtain a filtered motion vector map; performing motion enhancement based on spectral characteristics on the filtered motion vector map to obtain a spectrally enhanced motion vector; Applying motion amplification algorithm to the temporal spectral stream according to the spectral enhancement motion vector, and constructing the enhanced motion vector field to obtain the enhanced motion vector field; Step S3: extracting thermal texture features from the time series spectrum stream to obtain a thermal texture feature map; Performing thermal texture dynamic change analysis on the thermal texture feature map to obtain a thermal texture change map; According to the thermal texture change map, the thermal texture anomaly is located and characterized for the preprocessed thermal image sequence to obtain a thermal texture anomaly map; Step S4: performing multimodal feature fusion on the time series spectral stream, enhanced motion vector field and thermal texture anomaly map to obtain a fused feature vector; Predict and generate risk confidence for the fused feature vector to obtain a risk confidence heat map; Step S5: collecting and preprocessing the situational information of the inspection area according to the sensor network deployment plan to obtain a situational information data set; Extract risk area information from the risk confidence heat map to obtain associated risk area information; Based on the associated risk area information and situational information data set, situational awareness warning is carried out to obtain an intelligent warning report.
2. The fire safety inspection method based on machine vision recognition according to claim 1 is characterized in that: Step S1 includes the following steps: Step S11: deploying a multispectral camera in the inspection area and deploying a sensor network to obtain a camera deployment plan and a sensor network deployment plan; Step S12: performing high-speed acquisition of multispectral image sequences according to the camera deployment scheme to obtain an original multispectral image stream; Step S13: performing multispectral image preprocessing on the original multispectral image stream to obtain a corrected multispectral image sequence; Step S14: performing multispectral image channel alignment and fusion on the corrected multispectral image sequence to obtain a multispectral image stack; Step S15: integrating the time-series spectral information of the multispectral image stack to obtain a time-series spectral stream.
3. The fire safety inspection method based on machine vision recognition according to claim 1 is characterized in that: Step S2 includes the following steps: Step S21: performing multi-spectral channel motion estimation on the time series spectral stream to obtain a channel motion vector diagram; Step S22: performing motion vector noise suppression and screening on the channel motion vector map to obtain a filtered motion vector map; Step S23: performing motion enhancement based on spectral characteristics on the filtered motion vector map according to the time-series spectral stream to obtain a spectrally enhanced motion vector; Step S24: using the spectral enhancement motion vector as guidance information, applying a motion amplification algorithm to the temporal spectral stream to obtain an amplified image sequence; Step S25: constructing an enhanced motion vector field for the amplified image sequence according to the spectral enhanced motion vector to obtain an enhanced motion vector field.
4. The fire safety inspection method based on machine vision recognition according to claim 3 is characterized in that: Step S23 includes the following steps: Step S231: Calculate the spectral characteristic response graph of the time series spectral stream to obtain a spectral characteristic response graph set; Step S232: constructing a spectral similarity weight map according to the spectral feature response atlas to obtain a spectral similarity weight atlas; Step S233: fusing the filtered motion vector map with the spectral weight according to the spectral similarity weight atlas to obtain a spectral weighted motion vector map; Step S234: performing time domain motion vector smoothing and accumulation on the spectrally weighted motion vector map to obtain a time domain smoothed motion vector map; Step S235: performing motion vector enhancement according to the time-domain smoothed motion vector image and the filtered motion vector image to obtain a spectrally enhanced motion vector.
5. The fire safety inspection method based on machine vision recognition according to claim 1 is characterized in that: Step S3 includes the following steps: Step S31: extracting an infrared image sequence from the time-series spectral stream, and performing a reverse geometric transformation using an enhanced motion vector field to obtain an aligned infrared image sequence; Step S32: performing thermal image preprocessing on the aligned infrared image sequence to obtain a preprocessed thermal image sequence; Step S33: extracting thermal texture features from the preprocessed thermal image sequence to obtain a thermal texture feature map; Step S34: performing a thermal texture dynamic change analysis on the thermal texture feature map to obtain a thermal texture change map; Step S35: locating and characterizing thermal texture anomalies of the preprocessed thermal image sequence according to the thermal texture change map to obtain a thermal texture anomaly map.
6. The fire safety inspection method based on machine vision recognition according to claim 5 is characterized in that: Step S34 includes the following steps: Step S341: Calculate the texture difference feature map on the thermal texture feature map to obtain the texture difference feature map; Step S342: Establishing a motion amplitude characteristic map for the enhanced motion vector field to obtain a motion amplitude characteristic map; Step S343: fusing the texture difference feature map and the motion amplitude feature map to obtain a fused feature score map; Step S344: performing time series anomaly detection on the fused feature score graph to obtain an anomaly score graph; Step S345: Generate a thermal texture change map according to the abnormal score map to obtain a thermal texture change map.
7. The fire safety inspection method based on machine vision recognition according to claim 5 is characterized in that: Step S35 includes the following steps: Step S351: generating a binary anomaly mask for the thermal texture change map to obtain a binary anomaly mask; Step S352: Connect the binary anomaly masks by component labeling to obtain a connected component labeling graph; Step S353: extracting abnormal regions from the preprocessed thermal image sequence according to the connected component label graph to obtain an abnormal region image block set; Step S354: Calculate the regional temperature statistical characteristics of the abnormal region image block set to obtain a regional temperature characteristic table; Step S355: Calculate the regional texture features of the abnormal region image block set to obtain a regional texture feature table; Step S356: classifying and marking abnormal patterns according to the regional temperature feature table and the regional texture feature table to obtain a marked abnormal region map; Step S357: Generate a thermal texture anomaly map based on the marked abnormal area map, the regional temperature feature table, and the regional texture feature table to obtain a thermal texture anomaly map.
8. The fire safety inspection method based on machine vision recognition according to claim 1 is characterized in that: Step S4 includes the following steps: Step S41: aligning multimodal feature data of the time series spectral stream, the enhanced motion vector field, and the thermal texture anomaly map to obtain an aligned multimodal feature set; Step S42: extracting multimodal deep features from the aligned multimodal feature set to obtain multimodal deep features; Step S43: performing feature-level fusion on the multimodal deep features and performing feature interactive learning to obtain a fused feature vector; Step S44: predict and generate risk confidence for the fused feature vector to obtain a risk confidence heat map.
9. The fire safety inspection method based on machine vision recognition according to claim 1 is characterized in that: Step S5 includes the following steps: Step S51: collecting and preprocessing context information of the inspection area according to the sensor network deployment plan to obtain a context information data set; Step S52: extracting risk areas from the risk confidence heat map to obtain risk area data; performing spatial information matching on the risk area data and the context information data set, and correlating the risk areas to obtain correlated risk area information; Step S53: weighting the situational factors according to the associated risk area information to obtain situational factor weighted data; performing risk correction according to the situational factor weighted data to obtain situational corrected risk confidence; Step S54: performing warning level classification according to the scenario-corrected risk confidence level to obtain warning level classification information; dynamically adjusting the threshold of the warning level classification information according to the scenario information data set to obtain warning level information; Step S55: Generate and push an intelligent warning report according to the warning level information, the associated risk area information and the time series spectrum flow to obtain an intelligent warning report.
10. A fire safety inspection system based on machine vision recognition, characterized in that: Used to execute the fire safety inspection method based on machine vision recognition as claimed in claim 1, the fire safety inspection system based on machine vision recognition comprises: The multi-spectral dynamic perception module is used to deploy inspection equipment in the inspection area and obtain the camera deployment plan and the sensor network deployment plan; perform multi-spectral dynamic perception according to the camera deployment plan to obtain the time-series spectral flow; The micro-motion amplification module is used to perform multi-spectral channel motion estimation on the time-series spectral stream and perform motion vector noise suppression to obtain a filtered motion vector map; perform motion enhancement based on spectral characteristics on the filtered motion vector map to obtain a spectrally enhanced motion vector; apply a motion amplification algorithm to the time-series spectral stream according to the spectrally enhanced motion vector, and perform an enhanced motion vector field construction to obtain an enhanced motion vector field; The thermal texture anomaly analysis module is used to extract thermal texture features from the time series spectral stream to obtain a thermal texture feature map; perform thermal texture dynamic change analysis on the thermal texture feature map to obtain a thermal texture change map; locate and characterize thermal texture anomalies on the preprocessed thermal image sequence based on the thermal texture change map to obtain a thermal texture anomaly map; The multimodal feature fusion module is used to perform multimodal feature fusion on the time series spectral flow, enhanced motion vector field and thermal texture anomaly map to obtain the fused feature vector; the risk confidence level of the fused feature vector is predicted and generated to obtain the risk confidence heat map; The situational awareness and early warning module is used to collect and preprocess situational information of the inspection area according to the sensor network deployment plan to obtain a situational information data set; extract risk area information from the risk confidence heat map to obtain associated risk area information; perform situational awareness and early warning based on the associated risk area information and the situational information data set to obtain an intelligent early warning report.
Citation Information
Patent Citations
Unmanned aerial vehicle forest fire prevention recognition system and method based on machine vision
CN118839288A
Collection bin module and flow cytometry data collection system
CN118858113A