Fire early warning method and system based on artificial intelligence

By using composite optical fibers and multiple types of sensors to collect multimodal data in industrial scenarios, and combining three-dimensional coordinate models and dynamic time rule algorithms for spatiotemporal alignment, artificial intelligence models deployed in the cloud and at the edge are trained. This solves the accuracy and real-time problems of traditional fire early warning methods in extreme environments, and achieves high-precision and rapid fire early warning.

CN120913375APending Publication Date: 2025-11-07TANGSHAN XINYADE RIPPLE TUBE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511132208.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Traditional fire early warning methods cannot accurately monitor in dark, high humidity, and high electromagnetic interference environments. Furthermore, existing multimodal AI solutions struggle to maintain high accuracy and low false alarm rates in extreme industrial scenarios, lack the ability to trace the root cause of faults, and suffer from poor real-time performance or bloated models that are difficult to deploy.

Method used

Multimodal data is collected using composite optical fibers and multiple types of sensors. Spatiotemporal alignment is performed by combining a preset three-dimensional coordinate model and a dynamic time rule algorithm to train a first artificial intelligence model deployed in the cloud. A lightweight second artificial intelligence model is then deployed on an edge device for real-time analysis using knowledge distillation technology. Multidimensional feature vectors and masks are used to improve the ability to identify fire risks.

Benefits of technology

It enables high-precision identification and rapid response to early fires in complex industrial environments, reduces false alarm and missed alarm rates, improves the accuracy, real-time performance and reliability of fire early warning, and supports timely early warning and risk location of fires.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913375A_ABST
    Figure CN120913375A_ABST
Patent Text Reader

Abstract

The invention provides a fire early warning method and system based on artificial intelligence. The method comprises the following steps: arranging a composite optical fiber and a target sensor in a to-be-detected area to collect historical multi-modal data so as to construct a multi-modal data sequence; performing space-time alignment on all modal data in the multi-modal data sequence to construct a training data set; extracting a multi-dimensional feature vector based on the training data set, determining a mask based on a preset three-dimensional coordinate model, and training a first artificial intelligence model according to the multi-dimensional feature vector and the mask; obtaining a target signal according to the trained first artificial intelligence model, and training a second artificial intelligence model according to the target signal; acquiring real-time multi-modal data through a composite optical fiber and a target sensor, and if real-time temperature modal data in the real-time multi-modal data reaches a preset triggering condition, triggering the trained second artificial intelligence model and inputting the real-time multi-modal data into the second artificial intelligence model to obtain a target fire risk probability, and triggering a target early warning operation according to the target fire risk probability.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of fire safety protection, in particular to a fire warning method and system based on artificial intelligence. BACKGROUND

[0002] With the continuous expansion of urban infrastructure and industrial parks, places such as pipe corridors, underground pipe networks, cable bridge, chemical plants and energy storage power stations are facing higher and higher fire and thermal runaway risks. Traditional fire warning methods rely on single sensors or visible light / infrared imaging, which cannot accurately monitor in dark, high humidity, high electromagnetic interference environments. Threshold-based linkage control often cannot take effective measures in time after false alarm or missed alarm, lacks the ability to trace the root cause of the fault, and leads to time-consuming and labor-intensive post-mortem. Existing multi-modal AI solutions either rely on cloud computing leading to poor real-time performance, or have bloated models that are difficult to deploy to on-site edge devices, or lack structural area perception and physical constraints, making it difficult to maintain high accuracy and low false alarm rate in industrial extreme scenarios.

[0003] Therefore, the present application provides a fire warning method and system based on artificial intelligence to solve one of the above technical problems. SUMMARY

[0004] The purpose of the present application is to provide a fire warning method and system based on artificial intelligence, which can solve at least one of the above technical problems. The specific scheme is as follows: According to the specific embodiment of the present application, in a first aspect, the present application provides a fire warning method based on artificial intelligence, comprising: Laying a composite optical fiber and a target sensor in a to-be-detected area, and collecting historical multi-modal data corresponding to a preset time period to construct a multi-modal data sequence, the multi-modal data including temperature modal, strain modal, vibration modal, carbon dioxide concentration modal and voiceprint modal; Based on a preset three-dimensional coordinate model and a dynamic time rule algorithm, all modal data in the multi-modal data sequence are spatio-temporally aligned to construct a training data set; Extracting a multi-dimensional feature vector based on the training data set, determining a mask based on the preset three-dimensional coordinate model, and training a first artificial intelligence model according to the multi-dimensional feature vector and the mask, wherein the first artificial intelligence model is deployed in the cloud; Obtaining a target signal according to the trained first artificial intelligence model, and training a second artificial intelligence model according to the target signal, wherein the second artificial intelligence model is deployed in a target device in the to-be-detected area; Collecting real-time multi-modal data through the composite optical fiber and the target sensor, and determining whether the real-time temperature modal data in the real-time multi-modal data meets a preset trigger condition; If the preset triggering condition is reached, the trained second artificial intelligence model is triggered and the real-time multi-modal data is input into the trained second artificial intelligence model, a target fire risk probability is obtained, and a target early warning operation is triggered according to the target fire risk probability.

[0005] According to the specific embodiment of the present application, in a second aspect, the present application provides an artificial intelligence-based fire early warning system, comprising: A historical data acquisition unit is configured to arrange a composite optical fiber and a target sensor in a to-be-detected area, and acquire historical multi-modal data corresponding to a preset time period to construct a multi-modal data sequence, wherein the multi-modal data includes temperature modal data, strain modal data, vibration modal data, carbon dioxide concentration modal data, and voiceprint modal data. A training set construction unit is configured to perform time-space alignment on all modal data in the multi-modal data sequence based on a preset three-dimensional coordinate model and a dynamic time warping algorithm to construct a training data set. A training unit is configured to extract a multi-dimensional feature vector based on the training data set, determine a mask based on the preset three-dimensional coordinate model, and train a first artificial intelligence model according to the multi-dimensional feature vector and the mask, wherein the first artificial intelligence model is deployed in the cloud; acquire a target signal according to the trained first artificial intelligence model, and train a second artificial intelligence model according to the target signal, wherein the second artificial intelligence model is deployed in a target device in the to-be-detected area. An early warning unit is configured to acquire real-time multi-modal data through the composite optical fiber and the target sensor, determine whether real-time temperature modal data in the real-time multi-modal data reaches a preset triggering condition, trigger the trained second artificial intelligence model and input the real-time multi-modal data into the trained second artificial intelligence model if the preset triggering condition is reached, obtain a target fire risk probability, and trigger a target early warning operation according to the target fire risk probability.

[0006] Compared with the prior art, the above-mentioned scheme of the present application has at least the following beneficial effects: (1) By arranging a composite optical fiber and a plurality of types of target sensors in a to-be-detected area to acquire historical multi-modal data, multi-dimensional information such as temperature, strain, vibration, carbon dioxide concentration, and voiceprint can be comprehensively captured, compared with a traditional single-modal monitoring method, the ability to capture early weak abnormal signals of fire is greatly improved, and the risk of false negatives is effectively reduced. (2) The preset three-dimensional coordinate model and the dynamic time warping algorithm are used to perform time-space alignment on the multi-modal data, which not only solves the problem of time offset of different modal data caused by sampling frequency and response time delay, but also ensures accurate correspondence of data in the physical space through three-dimensional space mapping, so that the training data set constructed has high physical consistency and spatial accuracy, providing a high-quality data basis for subsequent model training, thereby improving the effectiveness and generalization ability of model training. (3) Based on the training data set, the multi-dimensional feature vector is extracted, and the mask is determined combined with the preset three-dimensional coordinate model, further mining the depth features and spatial structure information in the data, so that the first artificial intelligence model can learn the correlation between the multi-modal data and the spatial structure when training in the cloud, and enhance the model's ability to identify and predict fire risks. Further, through the knowledge distillation technology, the target signal is obtained according to the trained first artificial intelligence model to train the second artificial intelligence model, and it is deployed on the target device in the detection area, realizing the lightweight deployment of the model, while retaining strong prediction ability, greatly reducing the calculation complexity and model size, and being able to perform real-time low-latency inference in the detection area, reducing data transmission delay and improving the response speed of early warning; (4) In the real-time monitoring stage, first, whether the real-time temperature modal data meets the preset trigger condition is judged as a preliminary screening mechanism, which reduces unnecessary calculation while quickly identifying potential risks. When the condition is met, the second artificial intelligence model is triggered and the real-time multi-modal data is input, and the target fire risk probability is obtained by comprehensive analysis of multi-dimensional information. Compared with single temperature threshold judgment, this method can more accurately assess fire risks and effectively reduce false positives. According to the accurate risk probability, the corresponding target warning operation is triggered, realizing the intelligentization of the whole chain from data acquisition, processing, analysis to warning decision, which not only can issue early warning in the early stage of fire, but also can more accurately locate the risk position through comprehensive analysis of multi-modal data, providing strong support for subsequent emergency disposal, and significantly improving the accuracy, real-time performance and reliability of fire warning in industrial scenarios. BRIEF DESCRIPTION OF DRAWINGS

[0007] Figure 1 A flowchart of an artificial intelligence-based fire warning method is shown; Figure 2 A unit block diagram of an artificial intelligence-based fire warning system according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0008] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0009] It should be particularly noted that the symbols and / or numbers existing in the specification, if not marked in the drawings, are not drawing marks.

[0010] The embodiments provided in the present application are embodiments of an artificial intelligence-based fire warning method. The embodiments will be described below with reference to the drawings Figure 1The embodiments of the present application are described in detail.

[0011] Figure 1 A flowchart of an artificial intelligence-based fire warning method is shown. The method comprises the following steps: S101, laying composite optical fibers and target sensors in the area to be detected, and collecting historical multi-modal data corresponding to a preset time period to construct a multi-modal data sequence, the multi-modal data including temperature modal data, strain modal data, vibration modal data, carbon dioxide concentration modal data, and voiceprint modal data; S102, based on a preset three-dimensional coordinate model and a dynamic time rule algorithm, temporally and spatially aligning all modal data in the multi-modal data sequence to construct a training data set; S103, extracting a multi-dimensional feature vector based on the training data set, determining a mask based on the preset three-dimensional coordinate model, and training a first artificial intelligence model according to the multi-dimensional feature vector and the mask, wherein the first artificial intelligence model is deployed in the cloud; S104, obtaining a target signal according to the trained first artificial intelligence model, and training a second artificial intelligence model according to the target signal, wherein the second artificial intelligence model is deployed in a target device in the area to be detected; S105, collecting real-time multi-modal data through the composite optical fibers and the target sensors, and determining whether the real-time temperature modal data in the real-time multi-modal data meets a preset trigger condition; S106, if the preset trigger condition is met, triggering the trained second artificial intelligence model and inputting the real-time multi-modal data into the second artificial intelligence model, obtaining a target fire risk probability, and triggering a target warning operation according to the target fire risk probability.

[0012] In some embodiments, the artificial intelligence-based fire warning method is suitable for fire risk perception and automatic warning response in industrial scenarios, especially for underground pipe corridors, energy storage power stations, chemical plants, and other places with high requirements for early fire warning accuracy and response speed. That is, the area to be detected includes underground pipe corridors, energy storage power stations, chemical plants, and the like, which are set according to actual application requirements.

[0013] As a feasible embodiment, the artificial intelligence-based fire warning method first lays composite sensing optical fibers and target sensors in the area to be detected. The composite optical fibers integrate a Raman scattering unit for temperature perception, a Brillouin scattering unit for strain monitoring, and a fiber micro-strain unit for vibration detection. The target sensors include a high-frequency acoustic sensor for voiceprint collection and a carbon dioxide concentration sensor for gas detection. Within a preset time period, such as from January 1 to January 7 for consecutive 7 days, sample every minute, continuously collect corresponding temperature modal data, strain modal data, vibration modal data, carbon dioxide concentration modal data, and voiceprint modal data, and form a historical multi-modal data sequence. The preset time is set according to actual application requirements, which is not limited here.

[0014] Subsequently, based on the three-dimensional space model corresponding to the to-be-detected region and a dynamic time warping (DTW) algorithm, historical multi-modal data sequences are processed in space-time alignment. The three-dimensional space model can be constructed by a field CAD (Computer Aided Design) drawing of the to-be-detected region, and records three-dimensional coordinate information of all target sensor layout nodes, composite optical fiber layout paths, structural units and sensing paths. For example, time alignment takes temperature modal data as a reference, and by calculating a time cost matrix between the temperature modal data and other modal data, time offsets caused by different sampling frequencies and response time delays are corrected. At the same time, by combining coordinate data of a spatial source point of each modal data, all modal data are mapped to a unified spatial coordinate system through coordinate projection and interpolation, and a training sample set with time consistency and spatial correspondence is constructed.

[0015] As a specific embodiment, multi-dimensional feature vectors are extracted from the training data set after space-time alignment. The features include temperature gradient, strain rate of change, vibration frequency spectrum energy, CO2 concentration gradient and MFCC (Mel-Frequency Cepstral Coefficients) voiceprint coefficient, etc. On this basis, in combination with key structures identified in the preset three-dimensional coordinate model, a structure prior mask is generated for each feature vector position. The key structures include cable joints, flanges, valves, etc. The mask is obtained by analyzing the CAD drawing through a deep learning model, and represents whether each spatial point is located at a high-risk structure position, and is introduced into the training of the subsequent artificial intelligence model as an attention weight. Finally, based on the multi-dimensional feature vectors and the mask, a first artificial intelligence model deployed in the cloud is trained. The model can adopt a multi-head attention Transformer (a deep learning model architecture) structure, and has strong multi-modal information fusion and key position identification capabilities. After the training of the first artificial intelligence model is completed, the fire risk probability distribution, the attention weight matrix and the hidden layer feature activation value output by the first artificial intelligence model are obtained as target signals, and a second artificial intelligence model deployed on a target device is trained by knowledge distillation. The target device includes a Jetson Nano (an edge computing device) edge device. The second artificial intelligence model is a lightweight network structure, which is usually composed of a convolutional neural network (CNN) and a gated recurrent unit (GRU), and can realize low-delay inference in a power-constrained environment. The distillation process keeps the consistency of model decision-making, and at the same time compresses the parameter volume to the range that can be deployed on the edge.

[0016] As a specific embodiment, during the system deployment operation, real-time multi-modal data is collected by the composite optical fiber and the target sensor. In each sampling period, the system first determines whether the real-time temperature modal data meets the preset triggering condition, such as the second derivative of the temperature exceeding a certain threshold, indicating that there may be a local overheating or insulation breakdown trend. When the first-level triggering condition is met, the system immediately activates the second artificial intelligence model, obtains the real-time feature vectors of temperature, strain, vibration, voiceprint, carbon dioxide concentration, etc. at the current time according to the real-time multi-modal data, and inputs them into the second artificial intelligence model to obtain the current fire risk probability. Finally, according to the comparison of the fire risk probability and the set multi-level risk threshold, the corresponding warning measures are triggered, such as alarm prompt, automatic shutdown of equipment, start of water cooling or evacuation broadcast, etc., so as to realize efficient early fire response.

[0017] Through this method, an end-to-end, full-modal, multi-level intelligent fire warning mechanism can be realized, which combines the advantages of physical modeling and artificial intelligence methods, and significantly improves the accuracy and response time of early fire identification in complex industrial environments.

[0018] In some embodiments, the step S101 of laying composite optical fibers and target sensors in the to-be-detected area and collecting historical multi-modal data corresponding to a preset time period to construct a multi-modal data sequence comprises: laying composite optical fibers along the cable bridge or pipe gallery in the to-be-detected area, and installing target sensors at preset key nodes at the same time, wherein the composite optical fibers integrate Raman scattering units, Brillouin scattering units and vibration sensing units, and the target sensors include carbon dioxide concentration sensors and high-frequency acoustic sensors; collecting historical temperature modal data, historical strain modal data and historical vibration modal data in a preset time period based on the Raman scattering units, Brillouin scattering units and vibration sensing units, respectively, collecting historical carbon dioxide concentration modal data and historical voiceprint modal data in a preset time period based on the carbon dioxide concentration sensors and high-frequency acoustic sensors, respectively; performing denoising processing on the historical temperature modal data, historical strain modal data, historical vibration modal data, historical carbon dioxide concentration modal data and historical voiceprint modal data, respectively, to construct a multi-modal data sequence containing temperature modal data sequence, strain modal data sequence, vibration modal data sequence, carbon dioxide concentration modal data sequence and voiceprint modal data sequence.

[0019] In some embodiments, to achieve high-precision fire warning in an industrial scene, first, physical deployment is performed in the area to be detected, including laying composite sensing optical fibers along the cable bridge or underground pipe gallery path, and installing specific types of environmental sensors at key nodes for monitoring. The composite optical fiber integrates Raman scattering units, Brillouin scattering units, and vibration sensing units to realize the synchronous acquisition of temperature, strain, and vibration and other physical quantities. Key nodes include cable joints, battery modules, valve flanges, and other areas where faults are prone to occur. Target sensors including carbon dioxide concentration sensors and high-frequency acoustic sensors are installed to monitor gas leaks and acoustic anomalies that accompany fires.

[0020] As a feasible embodiment, after the physical deployment is completed, data acquisition tasks are performed within a preset time period. Specifically, historical temperature modal data within the preset time period is obtained by the Raman scattering unit, reflecting the spatial distribution and time trend of temperature changes in the environment. Historical strain modal data is obtained by the Brillouin scattering unit, which is used to perceive the deformation behavior of the material structure under thermal expansion and contraction or mechanical disturbance. Historical vibration modal data is obtained by the vibration sensing unit, which identifies high-frequency mechanical fluctuations and abnormal vibration patterns. On the other hand, CO2 concentration change data recorded by the carbon dioxide concentration sensor and acoustic modal data collected by the high-frequency acoustic sensor within the preset time period are obtained simultaneously, which are used to capture the release of chemical gases and characteristic acoustic signals produced in the early stages of a fire.

[0021] As a specific embodiment, after data acquisition is completed, the historical raw data of the above five modalities are respectively subjected to denoising processing to improve the reliability of subsequent analysis. The historical temperature modal data is removed from non-fire temperature rise noise caused by optical fiber interference or environmental baseline changes. The historical strain modal data and the historical vibration modal data are removed from periodic signals caused by background mechanical disturbance. The CO2 concentration modal data is filtered to remove transient disturbances caused by personnel activities or ventilation fluctuations. The acoustic modal data uses Mel-frequency cepstral coefficient extraction and filtering methods to separate high-frequency non-structural noise. Through the above preprocessing steps, clean, continuous, and stable five modal data sequences are obtained, including temperature modal data sequence, strain modal data sequence, vibration modal data sequence, carbon dioxide concentration modal data sequence, and acoustic modal data sequence.

[0022] In some embodiments, the preset three-dimensional coordinate model and the dynamic time rule algorithm are used to perform spatio-temporal alignment on all modal data in the multi-modal data sequence in step S102 to construct the training data set, including: constructing a three-dimensional coordinate system corresponding to the detection area, identifying the layout nodes of the target sensors and the layout path of the composite optical fiber in the three-dimensional coordinate system to obtain the preset three-dimensional coordinate model; determining the coordinate data corresponding to all modal data in the multi-modal data sequence based on the preset three-dimensional coordinate model; taking all coordinate data corresponding to the temperature modal data sequence in the multi-modal data sequence as a reference, calculating the distances of the temperature modal data sequence respectively with the strain modal data sequence, the vibration modal data sequence, the carbon dioxide concentration modal data sequence and the voiceprint modal data sequence at each time point to form a corresponding time cost matrix; based on the corresponding time cost matrix, the strain modal data sequence, the vibration modal data sequence, the carbon dioxide concentration modal data sequence and the voiceprint modal data sequence are respectively adjusted to be aligned with the temperature modal data sequence in time and space by the dynamic time rule algorithm, and the first coordinate data corresponding to all modal data in the multi-modal data sequence is updated; and based on the multi-modal data sequence after spatio-temporal alignment and the first coordinate data corresponding to all modal data therein, the training data set is constructed.

[0023] In some embodiments, to achieve the physical consistency and spatial correspondence of multi-modal data and support the accurate training of subsequent models, a three-dimensional coordinate system of the detection area is first constructed. The three-dimensional coordinate system is based on the actual spatial layout of the industrial field in the detection area, and is modeled in combination with spatial structures such as cable bridge, pipeline trend and equipment distribution. By introducing CAD drawing data and field layout information, the spatial calibration of the composite optical fiber layout path and the layout nodes of various target sensors is completed in the three-dimensional coordinate system, thereby forming a preset three-dimensional coordinate model covering the detection area, which is used to represent the positional relationship of multi-modal data in the real space.

[0024] As a feasible embodiment, after the establishment of the preset three-dimensional coordinate model, the multi-modal data collected in time series is assigned with spatial position information. Specifically, according to the preset three-dimensional coordinate model, the coordinate data of the corresponding spatial node when each modal data is collected is retrieved respectively, and is labeled as the coordinate attribute of the modal data to form a spatial annotation data set including temperature, strain, vibration, carbon dioxide concentration and voiceprint modal data. For example, the temperature modal data is derived from the composite optical fiber laid on the surface of the cable, and the voiceprint modal data is derived from the high-frequency acoustic sensor located on the side of the battery module. The original sampling points of the two in the three-dimensional space are separated from each other, and the spatial correlation needs to be established through the preset three-dimensional coordinate model.

[0025] As a specific embodiment, in order to solve the problem of time sequence offset caused by differences in sampling frequency, response delay or sensing mechanism of multi-modal data, the coordinate sequence of all temperature modal data in the temperature modal data sequence is used as a time reference benchmark to construct a multi-modal time alignment strategy. With this benchmark as the center, the Euclidean distance of each data in the temperature modal data sequence and each data in the strain modal data sequence, each data in the vibration modal data sequence, each data in the carbon dioxide concentration modal data sequence, and each data in the voiceprint modal data sequence at each time point is calculated to generate five time cost matrices. In each cost matrix, not only the difference in data acquisition time is considered, but also the physical response relationship between each modal is considered, and the elements of the cost matrix are weighted to ensure that the alignment priority of the key modal is higher. For example, the thermal expansion and contraction correlation between temperature and strain is preferentially emphasized. This weight design ensures that the alignment error between key modes such as temperature-strain is minimized. Based on the time cost matrix, the DTW algorithm is used to search for the optimal time alignment path, so that the cumulative alignment cost of different modal sequences is minimized, and the optimal time alignment path defines the time mapping relationship between modes. For example, the long time sequence window of the strain modal data sequence and the short window of the temperature modal data sequence are aligned through nonlinear mapping to compensate for the time offset caused by the difference in sampling frequency. According to the optimal time alignment path, the time offset Δt of each modal data sequence relative to the benchmark is calculated, for example, the strain modal data sequence at time j corresponds to the time i+Δt of the temperature modal data sequence. Set the time offset threshold (such as ±1ms), discard the data window whose offset exceeds the threshold, and finally adjust each modal data sequence to the unified time axis through linear interpolation to ensure that the aligned data meets the time sequence accuracy requirements of the detection area.

[0026] For example, in a certain city underground pipe gallery scene, the temperature modal data series is selected as the benchmark sequence, and the cost path is constructed for each 10ms of high-frequency modal data such as strain and vibration, and the optimal time alignment path is matched through the DTW algorithm. In this process, a pre-set three-dimensional coordinate model is used to constrain the corresponding relationship of each modal data in the spatial dimension. The time index of the non-benchmark modal data is dynamically adjusted in the above processing process to align it to the same time reference point in a physical sense, and its spatial position in the three-dimensional coordinate system is also updated, that is, the first coordinate data after alignment is generated.

[0027] Through the above spatio-temporal alignment, the temperature, strain, vibration, carbon dioxide concentration and voiceprint modal data at each time are structured and integrated according to the aligned time points and unified coordinates, and a training data set is constructed. Each data sample in the training data set is composed of multi-modal data at each time and its spatial position code. The training data set has the following characteristics: all modal data share a unified time label, have spatial coordinate consistency, and the physical dependence relationship between modalities has been implicitly expressed through time regularization. The training data set provides a standardized input for subsequent multi-modal feature fusion and artificial intelligence model training.

[0028] In addition, after alignment, a generative adversarial network with gradient penalty can be used to generate multiple rare failure samples under the physical constraint of ensuring that the temperature gradient is positively correlated with strain, to synthesize an abnormal multi-modal data set. The training data set and the synthesized abnormal multi-modal data set are labeled with positive and negative sample labels for the next step of training. This mechanism effectively enhances the discriminability of abnormal samples during model training, significantly improving the accuracy of fire warning.

[0029] In some embodiments, the step S103 of extracting a multi-dimensional feature vector based on the training data set and determining a mask based on a preset three-dimensional coordinate model comprises: for each training sample in the training data set, extracting a temperature feature vector, a strain feature vector, a vibration feature vector, a carbon dioxide concentration feature vector and a voiceprint feature vector to obtain a multi-dimensional feature vector including the temperature feature vector, the strain feature vector, the vibration feature vector, the carbon dioxide concentration feature vector and the voiceprint feature vector; identifying second coordinate data of the preset key structure in the preset three-dimensional coordinate model, determining a key area based on the second coordinate data, and determining a mask according to the key area.

[0030] In some embodiments, to achieve accurate identification and high-confidence modeling of potential fire risks in an industrial environment, a feature extraction operation is performed for each training sample in the constructed training dataset. Specifically, the multi-modal data contained in each training sample is sequentially subjected to structural analysis, and corresponding temperature feature vectors, strain feature vectors, vibration feature vectors, carbon dioxide concentration feature vectors, and voiceprint feature vectors are extracted. For the temperature modal data in the training sample, statistical quantities such as mean, standard deviation, and second derivative are calculated to form a temperature feature vector reflecting the temperature change trend and fluctuation characteristics. For the strain modal data in the training sample, frequency domain features are obtained using Fourier transform, and strain rate and other indicators are combined to form a strain feature vector. For the vibration modal data in the training sample, vibration feature vectors are constructed by extracting information such as energy distribution and peak frequency in a specific frequency band (e.g., 20-50 kHz). For the carbon dioxide concentration modal data in the training sample, carbon dioxide concentration feature vectors are obtained based on concentration rate and concentration gradient. For the voiceprint modal data in the training sample, voiceprint feature vectors are generated using MFCC and other methods. Combining these feature vectors, a multi-dimensional feature vector containing multi-modal information is obtained.

[0031] As a feasible embodiment, to achieve focused learning of the model on key structures, a preset three-dimensional coordinate model is further called for spatial analysis. In the preset three-dimensional coordinate model, the positions of preset key structures are marked by semantic recognition, including but not limited to cable joints, transformer terminals, flange interfaces, valve nodes, and other high-risk equipment parts, and the corresponding second coordinate data is determined according to the marking. By filtering these structure nodes with fault sensitivity in the preset three-dimensional coordinate model, the corresponding second coordinate data can be identified and defined as a key area. Centered on the key structure, a certain range (e.g., radius 0.5 meters) is set to determine the key area. Then, a binary mask matrix is generated according to the key area, in which the positions corresponding to the key area are assigned a value of 1 and the other areas are assigned a value of 0, thereby obtaining a mask for subsequent model training. The mask can guide the model to focus on the feature information of the key structure area, explicitly enhance the model's attention to the key structure features, and thus improve the model's ability to identify potential fire precursors such as electrical local overheating, sealing leakage, and vibration anomalies.

[0032] For example, in the application scenario of a certain chemical plant, one training sample in the training data set records the multi-modal data of the equipment running in a certain period of time. A multi-dimensional feature vector is extracted from the sample, in which the temperature feature vector shows that the temperature change rate is abnormally high at a certain time, the strain feature vector reflects that the strain fluctuation of the local position of the equipment is intensified, the vibration feature vector shows energy concentration phenomenon in a certain frequency band, the carbon dioxide concentration feature vector shows a slow rising trend, and the voiceprint feature vector captures abnormal high-frequency noise signals. At the same time, through the analysis of the preset three-dimensional coordinate model, the position coordinates of the valve in the region corresponding to the sample are identified, and a key region with a radius of 0.5 meters is determined around the valve to generate a mask. In the subsequent model training process, the mask gives higher weight to the multi-modal features of the valve region, effectively improving the recognition ability of potential fire risks such as valve leakage.

[0033] In some embodiments, the step S103 of training the first artificial intelligence model according to the multi-dimensional feature vector and the mask comprises: constructing a multi-modal fusion model with a structure prior mask mechanism as the first artificial intelligence model; reducing the multi-dimensional feature vector to a preset dimension and splicing to obtain a first fusion vector, inputting the first fusion vector and the mask into the first artificial intelligence model, obtaining a first training loss value according to a first output result of the first artificial intelligence model and a preset multi-component loss function; if the first training loss value is greater than a first loss threshold, adjusting a first training parameter of the first artificial intelligence model, and returning to the step of inputting the first fusion vector and the mask into the attention mechanism layer of the first artificial intelligence model until the first training loss value is less than the first loss threshold.

[0034] In some embodiments, to realize the deep fusion of multi-modal physical perception information and improve the perception ability of fire risks, a multi-modal fusion model with a structure prior mask mechanism is constructed as the first artificial intelligence model. The first artificial intelligence model adopts a Transformer architecture based on a multi-head attention mechanism, which can fuse multi-source heterogeneous information such as temperature, strain, vibration, voiceprint, and carbon dioxide concentration, and introduces a structure prior mask to enhance the attention allocation of the model to the key region. The structure prior mask is derived from the spatial annotation results of the key structure in the preset three-dimensional coordinate model, and after encoding processing, it realizes consistent weighting with the attention calculation module of the fusion model.

[0035] As a feasible embodiment, before model training, the multi-dimensional feature vectors extracted from the training data set are preprocessed, specifically including reducing the dimension of the feature vector corresponding to each modality to a unified preset dimension (such as 64 dimensions) to ensure the dimension consistency of each modality when splicing and fusing. The dimension reduction processing is completed by using 1x1 convolution or linear transformation, and then the five types of modal features are combined into a first fusion vector through splicing operation. For example, in the energy storage substation scene, the multi-dimensional feature vectors after dimension reduction are spliced into a unified 320-dimensional fusion vector. The first fusion vector and the structure prior mask generated before are jointly used as input and sent into the first artificial intelligence model.

[0036] As a specific embodiment, during the model training phase, after receiving the fusion vector and the mask, the first artificial intelligence model first embeds the mask into the multi-head attention mechanism layer, which is used to dynamically enhance the attention to the key area features in the attention weight calculation. The first artificial intelligence model then outputs a first prediction result (such as fire risk probability distribution, attention weight matrix, and hidden layer feature activation value, etc.). In an example, the multi-head attention mechanism layer adopts 8-head attention calculation, and its formula is: Attention( Q , K , V )=softmax( + E mask) V ; Wherein, the query matrix Q is mapped from the temperature feature vector, the key matrix K is derived from the strain feature vector, and the value matrix V integrates the temperature, strain, vibration, voiceprint, and carbon dioxide concentration, etc. The mask E mask is generated based on a preset three-dimensional coordinate model, and a weight bias is applied to the key structure area such as cable joint and valve, so as to force the model to focus on the feature correlation of high-risk position, is the dimension of the key matrix, generally equal to 128. This mechanism captures the different dimensional dependency relationships of multi-modal data through 8 parallel attention heads, and realizes feature enhancement of key areas combined with the mask mechanism, which improves the recognition ability of fire risk features.

[0037] As a specific embodiment, the preset multi-component loss function is composed of a cross-entropy loss, a mask regularization loss, and a feature smoothing loss. The cross-entropy loss measures the deviation of the model prediction probability from the true risk label, the mask regularization loss imposes a higher penalty on the prediction error of the key area, such as a 50% increase in the error weight of the cable joint area, and the feature smoothing loss ensures that the change of adjacent feature vectors conforms to the physical law, such as the continuity of the temperature gradient. The weights of the cross-entropy loss, the mask regularization loss, and the feature smoothing loss can be set according to the application, for example, the weight proportions can be 0.6, 0.3, and 0.1, respectively. The formula of the preset multi-component loss function is: ; wherein, is the first training loss value, is the cross-entropy loss, is the feature smoothing loss, is the mask regularization loss, a is the weight of the cross-entropy loss, b is the weight of the feature smoothing loss, and c is the weight of the mask regularization loss.

[0038] The formula of the cross-entropy loss is: ; wherein, is the probability value of the first artificial intelligence model predicting the sample as a fire risk, with a value range of [0, 1]; is a sample category weight coefficient, used to balance the quantity difference between "fire risk" and "normal sample" in fire warning; is a focusing parameter, used to control the weight attenuation of easy-to-classify samples (such as obvious non-fire areas), so that the model focuses more on difficult-to-classify fire risk samples.

[0039] The formula of the feature smoothing loss is: ; wherein, is the strain modal data predicted by the first artificial intelligence model, reflecting the physical deformation of the monitored area; and T is the collected temperature modal data.

[0040] The formula of the mask regularization loss is: ; wherein, is a mask matrix; is the fire risk probability predicted by the first artificial intelligence model, is the true fire situation.

[0041] As a specific embodiment, when the calculated first training loss value is greater than a preset first loss threshold (such as 0.8), the first training parameters of the first artificial intelligence model are adjusted, the first training parameters mainly including the transformation weights of the query, key and value of each attention head in the Transformer, the position encoding matrix and the weight bias of the full connection layer. The training process traces back to the attention mechanism layer when the fusion vector and the mask input the model, the attention distribution is recalculated and the gradient is updated. Through repeated optimization, the loss value is finally stable and lower than the first loss threshold, the final model converges and has strong cross-modal feature extraction capability and spatial attention capability, which can be used for subsequent risk judgment and knowledge distillation. This embodiment shows that the Transformer model combined with the structural prior mask mechanism can significantly improve the response accuracy of key area abnormalities while maintaining the ability to fuse complex multi-modal features, and the training stability and convergence efficiency are also significantly better than the traditional mask-free Transformer architecture.

[0042] In some embodiments, the step S104 of obtaining the target signal according to the trained first artificial intelligence model, and training the second artificial intelligence model according to the target signal, comprises: extracting the fire risk probability distribution, the attention weight matrix and the hidden layer feature activation value from the output layer, the attention mechanism layer and the hidden layer of the trained first artificial intelligence model respectively, and taking the fire risk probability distribution, the attention weight matrix and the hidden layer feature activation value as the target signal; constructing a neural network model including feature extraction, time series modeling and classification output as the second artificial intelligence model; inputting the target signal and the first fusion vector into the second artificial intelligence model, obtaining a second training loss value according to the second output result of the second artificial intelligence model and a preset distillation loss function; if the second training loss value is greater than a second loss threshold, adjusting the second training parameters of the second artificial intelligence model, and returning to the step of inputting the target signal and the first fusion vector into the second artificial intelligence model, until the second training loss value is less than the second loss threshold.

[0043] In some embodiments, in order to migrate the high-precision decision logic learned by the first artificial intelligence model in a complex environment to the second artificial intelligence model, the training process of the second artificial intelligence model is carried out based on the knowledge distillation mechanism. The core of this process is to extract structured target signals from the trained first artificial intelligence model as the target output to guide the training of the second artificial intelligence model, so as to improve its decision-making ability in a low-power environment.

[0044] As a feasible embodiment, first, three types of key output signals are extracted from the first artificial intelligence model after training, specifically including: the fire risk probability distribution given by the model output layer, which is the final prediction result of the fusion model when processing the training samples; the attention weight matrix output by the multi-head attention mechanism layer in the Transformer architecture, which reflects the attention degree of the model to different modal and spatial location features; and the feature activation value generated by the intermediate hidden layer, which is used to describe the intermediate representation learning state of the model to the input multi-modal features. These three types of information together constitute the so-called target signal.

[0045] As a specific embodiment, when constructing the second artificial intelligence model, a lightweight neural network structure integrating feature extraction, time series modeling and classification output is adopted, which is composed of a feature extraction module composed of three layers of convolutional neural network (CNN), a time series modeling module composed of a bidirectional gated recurrent unit (BiGRU), and a fire risk classification output module composed of a fully connected layer. The model has the advantages of small model size, fast inference speed and low parameter quantity, and is suitable for deployment in edge devices such as Jetson Nano with limited computing power. During training, the target signal extracted from the first artificial intelligence model and the previously constructed first fusion vector are used as training input together and sent into the second artificial intelligence model. The model outputs its corresponding second prediction result, and compares it with the fire risk probability distribution, attention weight and feature activation value in the target signal. According to the preset distillation loss function, the second training loss value is calculated. The preset distillation loss function usually includes multiple components, such as KL (relative entropy) divergence, hidden layer feature matching loss and cross-entropy loss. The KL divergence is used to measure the difference between the fire risk probability distribution output by the second artificial intelligence model and the fire risk probability distribution output by the first artificial intelligence model. The hidden layer feature matching loss is usually calculated by L2 norm, which is used to constrain the feature activation value of the intermediate layer of the second artificial intelligence model to be consistent with the feature activation value of the corresponding hidden layer of the first artificial intelligence model. The cross-entropy loss is used to measure the classification error of the second artificial intelligence model prediction result and the true fire label. The formula of the preset distillation loss function is: ; Where α, β, γ are weight adjustment factors, set according to the actual model training performance; is the second training loss value, is the KL divergence loss, is the hidden layer feature matching loss, is the cross-entropy loss.

[0046] At the beginning of the model training, the second training loss value is usually high, and the second training parameters of the second artificial intelligence model are adjusted according to the error feedback, including the convolution kernel weight, the GRU internal door control parameter and the output layer weight bias, and the training iteration is re-executed in the input stage. The process continues until the second training loss value is lower than the set second loss threshold, and the model training is completed. In actual deployment testing, the second artificial intelligence model has far lower inference delay and computational resource consumption than the first model while maintaining the approximate accuracy of the first model, successfully achieving the knowledge transfer and lightweight deployment goals.

[0047] In some embodiments, the judgment in step S105 whether the real-time temperature modal data in the real-time multi-modal data reaches the preset triggering condition comprises: calculating a temperature first derivative and a temperature second derivative based on the real-time temperature modal data, the previous time temperature modal data corresponding to the real-time temperature modal data, and the collection time interval; judging whether the real-time temperature modal data is greater than a preset temperature alarm threshold, whether the temperature first derivative is greater than a preset temperature change rate threshold, or whether the temperature second derivative is greater than a preset temperature change mutation threshold; if the real-time temperature modal data is greater than the preset temperature alarm threshold, the temperature first derivative is greater than the preset temperature change rate threshold, or the temperature second derivative is greater than the preset temperature change mutation threshold, it is determined that the real-time temperature modal data reaches the preset triggering condition.

[0048] In some embodiments, in the real-time monitoring process of fire warning, whether to trigger the second artificial intelligence model is judged based on the collected real-time temperature modal data. Specifically, the real-time temperature modal data is obtained, and the previous time temperature modal data corresponding to the real-time temperature modal data and the time interval for collecting the data of the two time points are called, and the temperature first derivative and the temperature second derivative are calculated using these information, wherein the temperature first derivative reflects the change rate of temperature with time, and the temperature second derivative reflects the change of the change rate of temperature. Then, the real-time temperature modal data is compared with the preset temperature alarm threshold (such as 60℃), the calculated temperature first derivative is compared with the preset temperature change rate threshold (such as 5℃ / minute), and the temperature second derivative is compared with the preset temperature change mutation threshold (such as 3℃ / minute2). As long as any one of the three comparisons meets the conditions of “the real-time temperature modal data is greater than the preset temperature alarm threshold”, “the temperature first derivative is greater than the preset temperature change rate threshold” and “the temperature second derivative is greater than the preset temperature change mutation threshold”, it is determined that the real-time temperature modal data reaches the preset triggering condition, and the subsequent second artificial intelligence model is triggered to perform fire risk assessment. That is, only when a significant temperature change trend is detected, the second artificial intelligence model is started to make prediction, which is beneficial to save computational resources. The preset temperature alarm threshold, the preset temperature change rate threshold and the preset temperature change mutation threshold are set according to the actual application scenario, which is not limited here.

[0049] In some embodiments, the triggering of the trained second artificial intelligence model and inputting the real-time multi-modal data into the trained second artificial intelligence model to obtain the target fire risk probability in step S106 comprises: splicing a second fusion vector after pre-processing the real-time multi-modal data; inputting the second fusion vector into the trained second artificial intelligence model to obtain a target output result containing the target fire risk probability.

[0050] In some embodiments, when it is necessary to use the trained second artificial intelligence model to perform fire risk assessment, the collected real-time multi-modal data will be pre-processed. The pre-processing process covers cleaning of each modality data such as temperature, strain, vibration, carbon dioxide concentration and voiceprint, such as removing abnormal jump values in temperature data, denoising voiceprint data and the like, and then splicing the processed each modality data according to the preset feature dimension and order to form a second fusion vector. The second fusion vector is input into the second artificial intelligence model which has completed training and is deployed on the edge device. The second artificial intelligence model performs operation and processing on the input vector according to the learned multi-modal feature correlation rule and fire risk judgment logic, and finally outputs a target output result containing a target fire risk probability. The result can intuitively reflect the possibility of fire occurrence in the current monitoring area, and provide a basis for subsequent early warning operations.

[0051] As a feasible embodiment, in the daily fire monitoring of a chemical industry park, real-time multi-modal data is collected by deployed composite optical fibers and various sensors. In an example, during pre-processing, for temperature modality data, a sliding average algorithm is used to remove error values caused by transient interference of equipment, and for voiceprint modality data, mel-frequency cepstral coefficients are used to extract effective features and filter background noise. After pre-processing each modality data, a second fusion vector is spliced according to the preset order of “temperature-strain-vibration-carbon dioxide concentration-voiceprint”, and the preset vector dimension is set according to the model input requirement, such as 128 dimensions. Then the second fusion vector is input into the trained second artificial intelligence model. The model internally performs rapid fusion analysis of multi-modal features through a lightweight neural network structure, and outputs a target fire risk probability between 0 and 1.

[0052] For example, in the fire warning scene of the urban underground pipe gallery, the real-time multi-modal data acquisition frequency is once per second. In the preprocessing stage, the Fourier transform is performed on the vibration modal data to screen out the high-frequency vibration components related to the initial characteristics of the fire, and the anomaly detection and correction are performed on the carbon dioxide concentration modal data to ensure the accuracy of the data. The preprocessed modal data is spliced into a second fusion vector, which is input into the second artificial intelligence model. The second artificial intelligence model is based on the knowledge learned by the first artificial intelligence model through knowledge distillation and combines the spatial structure prior of the pipe gallery scene to process the fusion vector and output the target fire risk probability. If the cable joint in the pipe gallery starts to heat due to poor contact, the risk probability output by the model will gradually increase, triggering the corresponding warning process.

[0053] For example, in the warehouse fire warning scene, in the preprocessing stage, the temperature modal data is normalized to map its value to the 0-1 interval, and the voiceprint modal data is truncated to the characteristic frequency band that may be generated by the fire. Then the second fusion vector is obtained by splicing, and the vector is input into the second artificial intelligence model. The model analyzes the input vector according to the learned change rule of multi-modal data when a fire occurs. When the experimental warehouse corner starts to slowly heat up due to the accumulation of flammable materials, and abnormal noise appears in the voiceprint modal data, the target fire risk probability output by the model gradually increases, accurately identifying the fire hazard, and verifying that the method can effectively utilize multi-modal data to predict the fire risk probability in the simulation scene.

[0054] In some embodiments, the step S106 of triggering a target warning operation according to the target fire risk probability comprises: determining whether the target fire risk probability is greater than a first risk level threshold and less than a second risk level threshold; if the target fire risk probability is greater than the first risk level threshold and less than the second risk level threshold, performing a sound and light alarm through a target device, and the triggered target warning operation is to mark the position of the risk area in the detected area; if the target fire risk probability is greater than the second risk level threshold, determining whether the target fire risk probability is greater than a third risk level threshold; if the target fire risk probability is greater than the second risk level threshold and less than the third risk level threshold, the triggered target warning operation is to cut off the power supply or close the local power distribution; if the target fire risk probability is greater than the third risk level threshold, the triggered target warning operation is to start the fire sprinkler system, turn on the water cooling system, or start the ventilation system.

[0055] In some embodiments, after obtaining the target fire risk probability output by the second artificial intelligence model, the risk level threshold system is preset for grading warning. By comparing the target fire risk probability with the first, second, and third risk level thresholds set in advance, different levels of warning operations are executed to achieve accurate disposal of fire risks. The first, second, and third risk level thresholds are set according to the actual application scene.

[0056] As a feasible embodiment, in a fire warning system of a large commercial complex, the first risk level threshold is set to 0.5, the second risk level threshold is set to 0.7, and the third risk level threshold is set to 0.9. When the target fire risk probability of a certain floor is monitored to be 0.6, which is greater than the first risk level threshold 0.5 and less than the second risk level threshold 0.7, the target device deployed in the region immediately starts the sound and light alarm device, emits a flashing red light and a buzzing alarm, and labels the risk area position in the detection area, and pushes the risk positioning information to the terminal device of the security personnel, prompting them to go to the investigation.

[0057] The application also provides a system embodiment for implementing the method steps of the above embodiments, based on the same explanation of the name meaning and the same technical effect as the above embodiments, which will not be described here.

[0058] As shown in Figure 2 The application provides an artificial intelligence-based fire warning system, which comprises: A historical data acquisition unit 1001 is configured to arrange a composite optical fiber and a target sensor in a detection area, and acquire historical multi-modal data corresponding to a preset time period to construct a multi-modal data sequence, wherein the multi-modal data includes temperature modal data, strain modal data, vibration modal data, carbon dioxide concentration modal data, and voiceprint modal data. A training set construction unit 1002 is configured to perform spatio-temporal alignment on all modal data in the multi-modal data sequence based on a preset three-dimensional coordinate model and a dynamic time warping algorithm, to construct a training data set. A training unit 1003 is configured to extract a multi-dimensional feature vector based on the training data set, determine a mask based on a preset three-dimensional coordinate model, and train a first artificial intelligence model according to the multi-dimensional feature vector and the mask, wherein the first artificial intelligence model is deployed in the cloud; acquire a target signal according to the trained first artificial intelligence model, and train a second artificial intelligence model according to the target signal, wherein the second artificial intelligence model is deployed in a target device in the detection area. An early warning unit 1004 is configured to acquire real-time multi-modal data through the composite optical fiber and the target sensor, and determine whether real-time temperature modal data in the real-time multi-modal data meets a preset trigger condition; if the preset trigger condition is met, the trained second artificial intelligence model is triggered and the real-time multi-modal data is input into the second artificial intelligence model, to obtain a target fire risk probability, and a target early warning operation is triggered according to the target fire risk probability.

[0059] Regarding the system in the above embodiments, the specific manner in which each module performs an operation has been described in detail in the embodiments related to the method, and will not be described in detail here.

[0060] Although the operations are described as being performed in a particular order in the figures, this should not be understood as requiring the operations to be performed in the particular order shown or in serial, or that all operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing can be advantageous.

[0061] The methods and systems of the application can be implemented using standard programming techniques, with rules-based logic or other logic that can be implemented in software, hardware, or both. Also, it should be appreciated that the words "module" and "system" are used herein to refer to a software, firmware, or hardware component that performs a particular function, and that the various modules and systems can be implemented using one or more hardware or software components.

[0062] Any of the steps, operations, or procedures described herein can be performed or implemented using one or more hardware or software modules, alone or in combination with other devices. In one embodiment, a software module is implemented using a computer program product comprising a computer-readable medium containing computer program code, which can be executed by a computer processor for performing any or all of the steps, operations, or procedures described herein.

[0063] The foregoing description of implementations of the application has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the application to the precise form disclosed, and various modifications and variations are possible in light of the above teachings or can be acquired from practice of the application. One skilled in the art will recognize that many possible variations can be made within the scope of the present application. The embodiments were chosen and described in order to best explain the principles of the application and its practical application, and to thereby enable others skilled in the art to best utilize the application.

[0064] Further, it should be understood that, throughout this disclosure, any reference to a method comprising two or more steps is intended to mean that the method can comprise the two or more steps in any suitable order, including sequentially or simultaneously.

[0065] Other embodiments of the application will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the application being indicated by the following claims.

[0066] It is to be understood that the application is not limited to the precise construction herein described and illustrated in the accompanying drawings, and that various modifications and changes can be made by those skilled in the art without departing from the scope thereof. The scope of the application is limited only by the claims appended hereto.

[0067] The above examples are only used to illustrate the technical solutions of the present application, but not limit it. Although the present application has been described in detail with reference to the foregoing embodiments, it should be understood by those skilled in the art that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalent ones. Such modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A fire warning method based on artificial intelligence, characterized in that, The method comprises the following steps: deploying a composite optical fiber and a target sensor in a to-be-detected area, and collecting historical multi-modal data corresponding to a preset time period to construct a multi-modal data sequence, wherein the multi-modal data includes temperature modal data, strain modal data, vibration modal data, carbon dioxide concentration modal data, and voiceprint modal data; spatially and temporally aligning all modal data in the multi-modal data sequence based on a preset three-dimensional coordinate model and a dynamic time rule algorithm to construct a training data set; extracting a multi-dimensional feature vector based on the training data set, determining a mask based on the preset three-dimensional coordinate model, and training a first artificial intelligence model according to the multi-dimensional feature vector and the mask, wherein the first artificial intelligence model is deployed in the cloud; obtaining a target signal according to the trained first artificial intelligence model, and training a second artificial intelligence model according to the target signal, wherein the second artificial intelligence model is deployed in a target device in the to-be-detected area; collecting real-time multi-modal data through the composite optical fiber and the target sensor, and determining whether real-time temperature modal data in the real-time multi-modal data reaches a preset trigger condition; if the preset trigger condition is reached, triggering the trained second artificial intelligence model and inputting the real-time multi-modal data into the second artificial intelligence model to obtain a target fire risk probability, and triggering a target warning operation according to the target fire risk probability.

2. The method of claim 1, wherein, The method comprises the following steps: deploying the composite optical fiber along a cable bridge or a pipe gallery in the to-be-detected area, and simultaneously installing the target sensor at a preset key node, wherein the composite optical fiber integrates a Raman scattering unit, a Brillouin scattering unit, and a vibration sensing unit, and the target sensor includes a carbon dioxide concentration sensor and a high-frequency acoustic sensor; collecting historical temperature modal data, historical strain modal data, and historical vibration modal data in the preset time period based on the Raman scattering unit, the Brillouin scattering unit, and the vibration sensing unit, respectively, and collecting historical carbon dioxide concentration modal data and historical voiceprint modal data in the preset time period based on the carbon dioxide concentration sensor and the high-frequency acoustic sensor, respectively; performing denoising processing on the historical temperature modal data, the historical strain modal data, the historical vibration modal data, the historical carbon dioxide concentration modal data, and the historical voiceprint modal data, respectively, to construct the multi-modal data sequence including a temperature modal data sequence, a strain modal data sequence, a vibration modal data sequence, a carbon dioxide concentration modal data sequence, and a voiceprint modal data sequence.

3. The method of claim 2, wherein, The method comprises the following steps: constructing a three-dimensional coordinate system corresponding to the to-be-detected area, identifying a deployment node of the target sensor and a deployment path of the composite optical fiber in the three-dimensional coordinate system, and obtaining the preset three-dimensional coordinate model; determining coordinate data corresponding to all modal data in the multi-modal data sequence based on the preset three-dimensional coordinate model; calculating distances between the temperature modal data sequence and the strain modal data sequence, the vibration modal data sequence, the carbon dioxide concentration modal data sequence and the voiceprint modal data sequence at each time point based on all coordinate data corresponding to the temperature modal data sequence in the multi-modal data sequence, to form a corresponding time cost matrix; aligning the strain modal data sequence, the vibration modal data sequence, the carbon dioxide concentration modal data sequence and the voiceprint modal data sequence in time and space to the temperature modal data sequence based on the corresponding time cost matrix by using the dynamic time warping algorithm, and updating the first coordinate data corresponding to all modal data in the multi-modal data sequence; constructing the training data set based on the multi-modal data sequence after time and space alignment and the first coordinate data corresponding to all modal data therein.

4. The method of claim 1, wherein, The extracting of the multi-dimensional feature vector based on the training data set and the determination of the mask based on the preset three-dimensional coordinate model include: extracting a temperature feature vector, a strain feature vector, a vibration feature vector, a carbon dioxide concentration feature vector and a voiceprint feature vector for each training sample in the training data set to obtain the multi-dimensional feature vector including the temperature feature vector, the strain feature vector, the vibration feature vector, the carbon dioxide concentration feature vector and the voiceprint feature vector; identifying second coordinate data of a preset key structure in the preset three-dimensional coordinate model, determining a key area based on the second coordinate data, and determining the mask according to the key area.

5. The method of claim 1, wherein, The training of the first artificial intelligence model according to the multi-dimensional feature vector and the mask includes: constructing a multi-modal fusion model with a structure prior mask mechanism as the first artificial intelligence model; pasting the first fusion vector obtained by reducing the dimension to a preset dimension to obtain a first fusion vector, inputting the first fusion vector and the mask into the attention mechanism layer of the first artificial intelligence model, and obtaining a first training loss value according to a first output result of the first artificial intelligence model and a preset multi-component loss function; if the first training loss value is greater than a first loss threshold, adjusting the first training parameter of the first artificial intelligence model, and returning to the step of inputting the first fusion vector and the mask into the attention mechanism layer of the first artificial intelligence model until the first training loss value is less than the first loss threshold.

6. The method of claim 5, wherein, The obtaining of the target signal according to the trained first artificial intelligence model and the training of the second artificial intelligence model according to the target signal include: extracting a fire risk probability distribution, an attention weight matrix and a hidden layer feature activation value from the output layer, the attention mechanism layer and the hidden layer of the trained first artificial intelligence model, respectively, and taking the fire risk probability distribution, the attention weight matrix and the hidden layer feature activation value as the target signal; constructing a neural network model including feature extraction, time series modeling and classification output as the second artificial intelligence model; inputting the target signal and the first fusion vector into the second artificial intelligence model, obtaining a second training loss value according to a second output result of the second artificial intelligence model and a preset distillation loss function; if the second training loss value is greater than a second loss threshold, adjusting a second training parameter of the second artificial intelligence model, and returning to the step of inputting the target signal and the first fusion vector into the second artificial intelligence model until the second training loss value is less than the second loss threshold.

7. The method of claim 1, wherein, The judgment whether the real-time temperature modal data in the real-time multi-modal data reaches a preset triggering condition comprises: calculating a temperature first-order derivative and a temperature second-order derivative based on the real-time temperature modal data, previous-time temperature modal data corresponding to the real-time temperature modal data, and a collection time interval; judging whether the real-time temperature modal data is greater than a preset temperature alarm threshold, whether the temperature first-order derivative is greater than a preset temperature change rate threshold, or whether the temperature second-order derivative is greater than a preset temperature change mutation threshold; if the real-time temperature modal data is greater than the preset temperature alarm threshold, the temperature first-order derivative is greater than the preset temperature change rate threshold, or the temperature second-order derivative is greater than the preset temperature change mutation threshold, it is determined that the real-time temperature modal data reaches the preset triggering condition.

8. The method of claim 1, wherein, The triggering of the trained second artificial intelligence model and the inputting of the real-time multi-modal data into the trained second artificial intelligence model to obtain a target fire risk probability comprises: splicing a second fusion vector after preprocessing the real-time multi-modal data; inputting the second fusion vector into the trained second artificial intelligence model to obtain a target output result containing the target fire risk probability.

9. The method of claim 1, wherein, The triggering of a target warning operation according to the target fire risk probability comprises: judging whether the target fire risk probability is greater than a first risk level threshold and less than a second risk level threshold; if the target fire risk probability is greater than the first risk level threshold and less than the second risk level threshold, performing an audible and visual alarm through the target device, and the triggered target warning operation is marking a risk area position in the to-be-detected region; if the target fire risk probability is greater than the second risk level threshold, judging whether the target fire risk probability is greater than a third risk level threshold; if the target fire risk probability is greater than the second risk level threshold and less than the third risk level threshold, the triggered target warning operation is cutting off power supply or closing local power distribution; if the target fire risk probability is greater than the third risk level threshold, the triggered target warning operation is starting a fire sprinkler system, turning on a water cooling system, or starting a ventilation system.

10. An artificial intelligence-based fire warning system, characterized by, The method comprises: a historical data acquisition unit configured to arrange a composite optical fiber and a target sensor in a to-be-detected region, and acquire historical multi-modal data corresponding to a preset time period to construct a multi-modal data sequence, the multi-modal data comprising temperature modal data, strain modal data, vibration modal data, carbon dioxide concentration modal data, and voiceprint modal data; a training set construction unit configured to perform time-space alignment on all modal data in the multi-modal data sequence based on a preset three-dimensional coordinate model and a dynamic time warping algorithm to construct a training data set; The training unit is configured to extract a multi-dimensional feature vector based on the training data set, determine a mask based on the preset three-dimensional coordinate model, train a first artificial intelligence model according to the multi-dimensional feature vector and the mask, wherein the first artificial intelligence model is deployed in the cloud; obtain a target signal according to the trained first artificial intelligence model, and train a second artificial intelligence model according to the target signal, wherein the second artificial intelligence model is deployed in a target device in the detection area to be detected; The early warning unit is configured to collect real-time multi-modal data through the composite optical fiber and the target sensor, and determine whether real-time temperature modal data in the real-time multi-modal data reaches a preset trigger condition; If the preset trigger condition is reached, the trained second artificial intelligence model is triggered and the real-time multi-modal data is input into the trained second artificial intelligence model, to obtain a target fire risk probability, and a target early warning operation is triggered according to the target fire risk probability.

Citation Information

Cited By

  • Energy storage power station fire prediction method based on spatio-temporal data

    CN121640679A