A Multimodal Energy Storage Power Station Fire Detection System Based on Image and Time Series Analysis
Through a multimodal detection system based on image and timing analysis, combined with YOLOv5 and LSTM models, the accuracy and response speed of fire detection in energy storage power stations are solved, and early identification and accurate early warning of fires are achieved.
Patent Information
- Application Number
- CN202411368343.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-29
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-09-29
AI Technical Summary
The existing fire detection methods in energy storage power plants have problems such as limited detection range, susceptibility to environmental factors, slow response speed, inaccurate feature extraction and lack of consideration for the dynamic characteristics of fires.
A multimodal detection system based on image and timing analysis is adopted, combined with the YOLOv5 object detection model and long-term short-term memory network LSTM, and accurate detection and early warning of flames and smoke are achieved through feature extraction, short-term fire decisions and long-term fire decision modules.
It improves the accuracy and response speed of fire detection, can identify fires in advance and reduce false alarm rates, and adapt to complex energy storage power station environments.
Smart Images

Figure QLYQS_1 
Figure QLYQS_3
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a multimodal energy storage power station fire detection system based on image and time series analysis. Background Art
[0002] With the rapid development of renewable energy, energy storage power stations are increasingly widely used in the power system; however, due to their high energy density and complex electrical equipment, energy storage power stations have certain fire risks; therefore, timely and accurately detecting and warning of fires is crucial for ensuring the safety of energy storage power stations;
[0003] Currently, traditional fire detection methods mainly rely on sensors such as smoke detectors and temperature detectors, but this method of relying on sensors for fire detection has the following disadvantages:
[0004] Limited detection range: The sensors can only cover a limited area and it is difficult to achieve comprehensive monitoring of the entire energy storage power station;
[0005] Prone to environmental factor interference: Smoke, temperature and other sensors are easily affected by environmental factors, resulting in false alarms or missed alarms;
[0006] Slow response speed: The sensors need a certain amount of time to send out alarm signals, making it difficult to achieve early warning;
[0007] To overcome the above limitations, in recent years, fire detection technologies based on video image analysis have gradually emerged, using computer vision technology to analyze video images, so as to achieve early detection and warning of fires; however, the existing video image analysis fire detection methods have the following disadvantages:
[0008] Inaccurate feature extraction: Existing methods mainly rely on manually designed features and it is difficult to effectively extract fire features, resulting in low detection accuracy;
[0009] Lack of consideration of fire dynamic features: Existing methods mainly focus on the static features of fires and lack consideration of fire dynamic features, while the occurrence and development of fires have time continuity, which is likely to lead to insufficient detection ability for early fires; Summary of the Invention
[0010] The purpose of the present invention is to solve the drawbacks existing in the prior art, and to propose a multimodal energy storage power station fire detection system based on image and time series analysis.
[0011] To achieve the above purpose, the present invention adopts the following technical solutions:
[0012] A multimodal energy storage power station fire detection system based on image and time series analysis, including a feature extraction module, a short-term fire decision-making module, and a long-term fire decision-making module;
[0013] The feature extraction module is used to extract the feature vectors of each frame of the video image;
[0014] The short-term fire decision-making module forms a time-series input sequence by arranging the feature vectors in chronological order, inputs the time-series input sequence into the long short-term memory network (LSTM), and preliminarily determines whether there is a fire in the video image according to the output result;
[0015] The long-term fire decision-making module divides the time-series input sequence into multiple time windows, judges whether there is a fire or not in each time window, votes on the output results of each time window, and finally determines whether there is a fire in the video image.
[0016] The present invention also proposes a multi-modal energy storage power station fire detection method based on image and time-series analysis, including the following sub-steps:
[0017] S1: Obtain the video image of the fire area and extract the feature vectors in each frame of the image;
[0018] Obtain the video image of the fire area, use the YOLOv5 object detection model to detect each frame of the obtained video image, and obtain the final feature vectors of each frame of the video image through the feature extraction module;
[0019] Including the following sub-steps:
[0020] S11: Obtain the video image and perform preprocessing;
[0021] Engineers obtain the video image of the fire area to be predicted in real time and perform preprocessing on each frame of the video image. The preprocessing includes scaling, normalization, etc.;
[0022] Specifically, use interpolation methods such as bilinear interpolation to scale each frame of the image to a specified size; normalize the pixel values of each frame of the image from the standard RGB range of 0-255 to 0-1 through the pixel normalization formula. The pixel normalization formula is as follows:
[0023] ;
[0024] Among them, is the normalized pixel value, and I is the original pixel value;
[0025] S12: Input the image into the target prediction model to generate the feature map corresponding to the image;
[0026] Each frame of the pre - processed image in step S11 is sequentially input into the YOLOv5 object detection model. The YOLOv5 object detection model processes the image through a feature extractor using convolutional layers and activation functions to generate a feature map corresponding to the image;
[0027] The feature map includes width, height, and number of channels;
[0028] The YOLOv5 object detection model has been pre - trained for flames and smoke;
[0029] S13: Generate a corresponding weight map according to the feature map;
[0030] The YOLOv5 object detection model divides the feature map into multiple grid cells and predicts A bounding boxes in each grid cell, where A is a positive integer;
[0031] The YOLOv5 object detection model predicts the class to which the region in the bounding box belongs and the confidence score under this class, that is, the bounding box contains coordinates and the corresponding confidence score; The classes include flames and smoke, and the confidence score corresponds to the possibility that the bounding box contains flames or smoke, and the confidence score is between 0 and 1;
[0032] According to the confidence scores of each bounding box included in the feature map, a weight map of the same size as the feature map is generated; Each pixel value of the weight map corresponds to the confidence score and the class to which each grid cell in the feature map belongs;
[0033] S14: Obtain a weighted feature map;
[0034] Create a weighted feature map of the same size as the feature map, with an initial value of 0;
[0035] Traverse each predicted bounding box, determine its corresponding position in the feature map according to the coordinates of the bounding box, and obtain the confidence score of the bounding box; Multiply the confidence score of the bounding box by the pixel value of the corresponding region in the feature map and accumulate it to the corresponding position of the weighted feature map; Realize the update of the initial weighted feature map;
[0036] S15: Obtain a feature vector;
[0037] Perform weighted average pooling on the weighted feature map through the following formula to obtain a feature vector; The formula is as follows:
[0038] ;
[0039] Where, is the feature vector of the i - th channel in the weighted feature map, refers to all pixel positions within the bounding box inside ; refers to the position coordinates in the weighted feature map at the -th channel eigenvalue, refers to the confidence score of the detection box ;
[0040] And use the normalization factor method to normalize it to obtain the final feature vector;
[0041] Loop the above steps until the final feature vector is obtained for each frame image in the video image.
[0042] S2: Make short-term fire decisions;
[0043] Set a threshold L in advance in combination with actual requirements;
[0044] Arrange the feature vectors of each frame image obtained in step S1 in chronological order, and the feature vectors of multiple consecutive frames form a time series input sequence;
[0045] The short-term fire decision module makes short-term fire decisions on the time series input sequence;
[0046] Specifically, input the feature vectors in the time series input sequence into the trained long short-term memory network LSTM in sequence, and the long short-term memory network LSTM outputs the fire probability value; if the fire probability value output by the long short-term memory network LSTM is greater than the threshold L, it is initially determined that there is a fire in this video image, otherwise not.
[0047] S3: Make long-term fire decisions;
[0048] The long-term fire decision module makes long-term fire decisions on the time series input sequence;
[0049] Specifically, divide the time input sequence into multiple small time windows, and make fire and non-fire judgments within each time window; count the number of fire decisions and non-fire decisions made by the long short-term memory network LSTM in all time windows; and obtain the long-term fire decision through the majority voting method; that is, when the sum of the number of fire decisions in all time windows is greater than the sum of the number of non-fire decisions, it is finally determined that there is a fire in this video image, otherwise not.
[0050] Compared with the prior art, the beneficial effects of the present invention are:
[0051] The method proposed by the present invention combines image feature extraction and time series analysis technologies. Through multi-modal fusion, it realizes early detection and warning of energy storage power station fires. The YOLOv5 object detection model is used to detect flames and smoke in each frame of the image, and a weighted feature map is generated. Weighted average pooling is performed on the weighted feature map to obtain a feature vector, which improves the accuracy of feature extraction. It solves the problem in the prior art that it is difficult to effectively extract fire features by relying on manually designed features, resulting in low detection accuracy.
[0052] At the same time, the long short-term memory network LSTM is used to analyze the dynamic behavior of consecutive video frames, make an accurate fire judgment in a short time, divide the time input sequence into multiple small time windows, vote on the output results of the time windows, and make a final fire decision. Combining short-term and long-term fire decisions significantly improves the accuracy and response speed of fire detection. Detailed implementation manners
[0053] To further understand the purpose, structure, features, and functions of the present invention, the following is a detailed description in conjunction with embodiments.
[0054] A multi-modal energy storage power station fire detection system based on image and time series analysis includes a feature extraction module, a short-term fire decision module, and a long-term fire decision module.
[0055] The feature extraction module is used to extract the feature vector of each frame of the video image.
[0056] The short-term fire decision module arranges the feature vectors in chronological order to form a time series input sequence, inputs the time series input sequence into the long short-term memory network LSTM, and initially determines whether there is a fire in the video image according to the output result.
[0057] The long-term fire decision module divides the time series input sequence into multiple time windows, makes fire and non-fire judgments within each time window, votes on the output results of each time window, and finally determines whether there is a fire in the video image.
[0058] The present invention also proposes a multi-modal energy storage power station fire detection method based on image and time series analysis, including the following sub-steps:
[0059] S1: Obtain the video image of the fire area and extract the feature vector of each frame of the image.
[0060] Obtain the video image of the fire area, use the YOLOv5 object detection model to detect each frame of the obtained video image, and obtain the final feature vector of each frame of the video image through the feature extraction module.
[0061] It includes the following sub-steps:
[0062] S11: Obtain video images and perform preprocessing;
[0063] Engineers obtain video images of the fire area to be predicted in real time and perform preprocessing on each frame of the video images. The preprocessing includes scaling, normalization, etc.;
[0064] Specifically, use interpolation methods such as bilinear interpolation to scale each frame of the image to a specified size; normalize the pixel values of each frame of the image from the standard RGB range of 0 - 255 to 0 - 1 through the pixel normalization formula. The pixel normalization formula is as follows:
[0065] ;
[0066] where, is the normalized pixel value, and I is the original pixel value;
[0067] By performing preprocessing such as scaling and normalization on the video images, the video images meet the size requirements of the YOLOv5 object detection model, avoiding errors caused by size mismatches. Through normalization, pixel values in different ranges can be adjusted to a relatively consistent range, which helps to accelerate the convergence speed of the model and improve the subsequent training efficiency.
[0068] S12: Input the image into the object prediction model to generate a feature map corresponding to the image;
[0069] Input each frame of the preprocessed image in step S11 into the YOLOv5 object detection model in sequence. The YOLOv5 object detection model processes the image through a feature extractor using convolutional layers and activation functions to generate a feature map corresponding to the image;
[0070] The feature map includes width, height, and the number of channels;
[0071] The YOLOv5 object detection model has been pre-trained for flames and smoke;
[0072] S13: Generate a corresponding weight map according to the feature map;
[0073] The YOLOv5 object detection model divides the feature map into multiple grid cells and predicts A bounding boxes in each grid cell, where A is a positive integer;
[0074] The YOLOv5 object detection model predicts the class to which the region in the bounding box belongs and the confidence score for that class, that is, the bounding box contains coordinates and corresponding confidence scores; the classes include fire and smoke, and the confidence score corresponds to the likelihood that the bounding box contains fire or smoke, with the confidence score ranging from 0 to 1;
[0075] Generate a weight map of the same size as the feature map based on the confidence scores of each bounding box contained in the feature map; each pixel value of the weight map corresponds to the confidence score and the class to which each grid cell in the feature map belongs;
[0076] S14: Obtain the weighted feature map;
[0077] Create a weighted feature map of the same size as the feature map, with an initial value of 0;
[0078] Traverse each predicted bounding box, determine its corresponding position in the feature map according to the coordinates of the bounding box, and obtain the confidence score of the bounding box; multiply the confidence score of the bounding box by the pixel value of the corresponding region in the feature map and accumulate it to the corresponding position in the weighted feature map; realize the update of the initial weighted feature map;
[0079] S15: Obtain the feature vector;
[0080] Perform weighted average pooling on the weighted feature map through the following formula to obtain the feature vector; the formula is as follows:
[0081] ;
[0082] where, is the feature vector of the i-th channel in the weighted feature map, refers to all pixel positions within the bounding box ; refers to the feature value of the -th channel at the position coordinate in the weighted feature map, refers to the confidence score of the detection box ;
[0083] And normalize it using the normalization factor method to obtain the final feature vector;
[0084] Loop the above steps until the final feature vector is obtained for each frame of the video image.
[0085] Detect fire and smoke in the video frame through the YOLOv5 object detection model and generate a weighted feature map to provide accurate spatial features for subsequent temporal analysis.
[0086] S2: Make short-term fire decisions;
[0087] Set a threshold L in advance according to the actual needs;
[0088] Arrange the feature vectors of each frame image obtained in step S1 in chronological order, and the feature vectors of multiple consecutive frames form a time series input sequence;
[0089] The short-term fire decision module makes a short-term fire decision on the time series input sequence;
[0090] Specifically, input the feature vectors in the time series input sequence into the trained long short-term memory network LSTM in order, and the long short-term memory network LSTM outputs a fire probability value; if the fire probability value output by the long short-term memory network LSTM is greater than the threshold L, it is preliminarily determined that there is a fire in this video image, otherwise there is no fire.
[0091] Furthermore, the training method of the long short-term memory network LSTM is as follows:
[0092] Collect video clips of various fire scenes as fire videos, and video clips containing confusing elements as non-fire videos, where the confusing elements include smoke, water vapor, clouds, lights, etc.;
[0093] Set n consecutive frame video clips as a sample, where n is a positive integer;
[0094] Extract corresponding numbers of positive samples and negative samples from the fire videos and non-fire videos respectively, and each sample contains n consecutive frame video clips;
[0095] Divide the samples into three parts: training, validation, and testing, to obtain a training set, a validation set, and a testing set; use the training set to train the long short-term memory network LSTM, verify the fire detection performance of the long short-term memory network LSTM through the validation set, and finally evaluate the accuracy of the fire detection of the long short-term memory network LSTM through the testing set; optimize the long short-term memory network LSTM according to the evaluation results, and the training is completed.
[0096] In the present invention, by collecting video clips containing confusing elements as non-fire videos as negative samples and combining them with positive samples to train the long short-term memory network LSTM, the judgment result of the long short-term memory network LSTM for the time input sequence is not easily affected by non-fire factors, solving the problem in the prior art that non-fire factors (such as smoke, water vapor, etc.) are easily misjudged as fires and the false alarm rate is high.
[0097] S3: Make a long-term fire decision;
[0098] The long-term fire decision module makes a long-term fire decision on the time series input sequence;
[0099] Specifically, the time input sequence is divided into multiple small time windows, and fire and non-fire judgments are made within each time window; by counting the number of fire decisions and non-fire decisions made by the long short-term memory network (LSTM) within all time windows; and through the majority voting method, a long-term fire decision is obtained; that is, when the sum of the number of fire decisions within all time windows is greater than the sum of the number of non-fire decisions, it is finally determined that there is a fire in this video image, otherwise there is no fire.
[0100] The present invention has been described by the above related embodiments. However, the above embodiments are only examples for implementing the present invention. It must be pointed out that the disclosed embodiments do not limit the scope of the present invention. On the contrary, modifications and refinements made without departing from the spirit and scope of the present invention fall within the scope of patent protection of the present invention.
Claims
1. A multi-modal energy storage power station fire detection method based on image and time series analysis, characterized in that: It includes the following steps: S1: Obtain the video image of the fire area and extract the feature vectors in each frame of the image; Obtain the video image of the fire area, use the YOLOv5 object detection model to detect each frame of the obtained video image, and obtain the final feature vector for each frame of the video image through the feature extraction module; It includes the following sub-steps: S11: Obtain the video image and perform preprocessing; S12: Input the image into the target prediction model to generate the corresponding feature map of the image; Input each frame of the preprocessed image in step S11 into the YOLOv5 object detection model in sequence. The YOLOv5 object detection model processes the image through the feature extractor using convolutional layers and activation functions to generate the corresponding feature map of the image; The feature map includes width, height, and number of channels; The YOLOv5 object detection model has been pre-trained for flames and smoke; S13: Generate the corresponding weight map according to the feature map; S14: Obtain the weighted feature map; Create a weighted feature map with the same size as the feature map, and the initial value is 0; Traverse each predicted bounding box, determine its corresponding position in the feature map according to the coordinates of the bounding box, and obtain the confidence score of the bounding box; Multiply the confidence score of the bounding box by the pixel value in the corresponding area of the feature map and accumulate it to the corresponding position of the weighted feature map; Realize the update of the initial weighted feature map; S15: Obtain the feature vector; S2: Make short-term fire decisions; Set a threshold L in advance according to the actual needs; Arrange the feature vectors of each frame of the image obtained in step S1 in chronological order, and the feature vectors of multiple consecutive frames form a time-series input sequence; The short-term fire decision module makes short-term fire decisions on the time-series input sequence; Specifically, input the feature vectors in the time-series input sequence into the trained long short-term memory network LSTM in sequence, and the long short-term memory network LSTM outputs the fire probability value; If the fire probability value output by the long short-term memory network LSTM is greater than the threshold L, it is initially determined that there is a fire in this video image, otherwise there is no fire; S3: Make long-term fire decisions; The long-term fire decision module makes long-term fire decisions on the time-series input sequence; Specifically, divide the time input sequence into multiple small time windows, and make fire and non-fire judgments within each time window; Count the number of fire decisions and non-fire decisions made by the long short-term memory network LSTM within all time windows; And through the majority voting method, obtain the long-term fire decision; That is, when the sum of the number of fire decisions within all time windows is greater than the sum of the number of non-fire decisions, it is finally determined that there is a fire in this video image, otherwise there is no fire.
2. A multi-modal energy storage power station fire detection method based on image and time-series analysis according to claim 1, characterized in that: In step S11, an engineer obtains the video image of the fire area to be predicted in real time and preprocesses each frame of the video image, and the preprocessing includes scaling and normalization; Specifically, use interpolation methods such as bilinear interpolation to scale each frame of the image to a specified size; normalize the pixel values of each frame of the image from the standard RGB range of 0 - 255 to 0 - 1 through the pixel normalization formula, and the pixel normalization formula is as follows: ; Among them, is the normalized pixel value, and I is the original pixel value.
3. A multi-modal energy storage power station fire detection method based on image and timing analysis according to claim 1, characterized in that: In step S13, the YOLOv5 object detection model divides the feature map into multiple grid cells, and predicts A bounding boxes in each grid cell, where A is a positive integer; The YOLOv5 object detection model predicts the category to which the region in the bounding box belongs and the confidence score under this category, that is, the bounding box contains coordinates and the corresponding confidence score; the categories include fire and smoke, and the confidence score corresponds to the possibility that the bounding box contains fire or smoke, and the confidence score is between 0 and 1; Generate a weight map of the same size as the feature map according to the confidence score of each bounding box included in the feature map; each pixel value of the weight map corresponds to the confidence score and the category to which each grid cell in the feature map belongs.
4. A multi-modal energy storage power station fire detection method based on image and timing analysis according to claim 1, characterized in that: In step S15, perform weighted average pooling on the weighted feature map through the following formula to obtain a feature vector; the formula is as follows: ; Among them, is the feature vector of the i-th channel in the weighted feature map, refers to all pixel positions within the bounding box ; refers to the feature value of the -th channel at the position coordinates in the weighted feature map, refers to the confidence score of the detection box ; And use the normalization factor method to normalize to obtain the final feature vector; Loop the above steps until the final feature vector is obtained for each frame of the video image.
5. A multi-modal energy storage power station fire detection method based on image and timing analysis according to claim 1, characterized in that: The training method of the long short-term memory network LSTM is as follows: Collect video clips of various fire scenes as fire videos, and video clips containing confusing elements as non-fire videos; Set n consecutive frames of video clips as a sample, where n is a positive integer; Extract corresponding numbers of positive and negative samples from the fire videos and non-fire videos respectively, and each sample contains n consecutive frames of video clips; Divide the samples into three parts: training, validation, and testing, to obtain a training set, a validation set, and a testing set; use the training set to train the long short-term memory network LSTM, verify the fire detection performance of the long short-term memory network LSTM through the validation set, and finally evaluate the accuracy of the fire detection of the long short-term memory network LSTM through the testing set; optimize the long short-term memory network LSTM according to the evaluation results, and the training is completed.
6. A multimodal energy storage power station fire detection system based on image and time series analysis, which is used to implement a multimodal energy storage power station fire detection method according to any one of claims 1-5, and is characterized in that: Including a feature extraction module, a short-term fire decision module, and a long-term fire decision module; The feature extraction module is used to extract the feature vector of each frame of the video image; The short-term fire decision module arranges the feature vectors in chronological order to form a time series input sequence, inputs the time input sequence into the long short-term memory network LSTM, and initially determines whether there is a fire in the video image according to the output result; The long-term fire decision-making module divides the time input sequence into multiple time windows, makes fire and non-fire judgments within each time window, votes on the output results of each time window, and finally determines whether there is a fire in the video image.
Citation Information
Patent Citations
Artificial intelligence-based autonomous alert system for real time remote fire and smoke detection in live video streams
US20240096187A1
Target detection method and apparatus
WO2021254205A1