Tunnel fire prediction method and system based on BiDS-Informer model

By introducing the BiDS-Informer model's bidirectional downsampling module, time-coded subtraction operation and dynamic attention mechanism in the Informer model, the Informer model's insufficient feature extraction capability, slow operation speed and inflexible attention mechanism in tunnel fire prediction are solved, and more efficient and accurate tunnel fire prediction is achieved.

CN120219911APending Publication Date: 2025-06-27NANJING FORESTRY UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510225120.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The application of Informer model in the field of tunnel fire prediction has problems such as insufficient feature extraction capabilities of sampling modules, slow operation speed, and the inability to flexibly adapt to different data types.

Method used

The BiDS-Informer model is adopted to enhance feature extraction capabilities by inserting a bidirectional downsampling module into the encoder and decoder; the encoder structure is optimized by time-encoding subtraction (TES) operation; a dynamic attention mechanism is introduced, and attention strategy and weight distribution are dynamically adjusted according to the input features.

Benefits of technology

It improves the model's deep correlation capture capability of complex data, improves processing speed and computing efficiency, and enhances prediction accuracy and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219911A_ABST
    Figure CN120219911A_ABST
Patent Text Reader

Abstract

The invention discloses a tunnel fire prediction method and system based on a BiDS-Informer model. The method comprises the following steps: step 1, collecting fire time sequence data of a tunnel; 2, a tunnel fire time sequence prediction model is constructed, and a BiDS-Informer model is selected as the tunnel fire time sequence prediction model and comprises an encoder and a decoder of the Informer model; the encoder input data is preprocessed tunnel fire time sequence data, and extracting features of the input data; the decoder converts output data of the encoder into a tunnel fire prediction result; downsampling modules are inserted between the multi-layer structures of the encoder and between the multi-layer structures of the decoder; step 3, training the BiDS-Informer model, and adjusting model parameters by using an optimization algorithm; and step 4, using the optimized BiDS-Informer model to predict the input image data in the original tunnel, and obtaining a final tunnel fire prediction result. The model provided by the invention provides important technical support for improving the reliability of the tunnel fire early warning system, and has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for predicting tunnel fires and belongs to the technical field of tunnel safety. Background Art

[0002] In recent years, the construction of tunnels has enhanced the flexibility of modern cities, especially in crossing natural barriers and connecting scattered areas, promoting regional economic development and the convenience of residents' lives. However, the high-intensity use of tunnels makes any minor oversight likely to lead to catastrophic consequences. The causes of tunnel fires are diverse, including electrical failures, mechanical equipment failures, and traffic accidents, etc. Therefore, accurately and timely predicting the tunnel fire situation is crucial for ensuring tunnel fire safety.

[0003] In the era of big data, machine learning technology has developed rapidly and is widely applied in various fields. The progress of machine learning highly depends on data, especially time series data. Time series data is data arranged in chronological order and has significant time correlation. Common examples include meteorological data, traffic flow, and environmental monitoring data, etc. To predict the future trend of time series data, researchers have proposed various methods, mainly divided into linear and non-linear prediction methods. Linear prediction methods are early classical methods, such as exponential smoothing method and autoregressive integrated moving average model, etc. Although these methods are suitable for short-term prediction, they have deficiencies in capturing non-linear relationships such as meteorological data. To overcome the limitations of linear methods, researchers have introduced non-linear models, such as BP neural network, support vector machine (SVM), RNN, and generative adversarial network (GAN), etc. These non-linear prediction methods can capture more comprehensive complex non-linear relationships in financial data, thus obtaining more accurate prediction results.

[0004] Driven by the latest progress in the fields of artificial intelligence and high-performance computing, numerous studies have proposed AI-based fire situation deduction methods. In forest fires, Thach respectively used SVM, random forest (RF), and multi-layer perceptron neural network (MLP-Net) to predict the fire situation. In various construction environments, Su proposed a data-driven fire situation prediction model based on convolutional neural network (CNN). In addition, Xie proposed an integrated model based on LSTM-TCN to predict the fire spread rate. However, LSTM (long short-term memory neural network) may lose information when processing long time series data, and in multi-step prediction, due to relying on the prediction of the previous moment, it is easy to accumulate errors and affect the prediction accuracy. As the dimension and volume of fire time series data increase, the advantages of the Transformer model in dealing with long sequence dependence gradually emerge. Miao proposed a time series prediction model based on Transformer, integrating various influencing factors, and made a more accurate prediction for the forest fire situation in a small area.

[0005] However, the self-attention mechanism of the Transformer model has the characteristics of quadratic time complexity and high memory usage, increasing the deployment cost and limiting its applications. To overcome these limitations, Zhou proposed a model for efficient long sequence time series forecasting (LSTF) based on an improved Transformer, called Informer. Informer effectively replaces the traditional self-attention with a dilated self-attention mechanism, achieving a time complexity of. In addition, Informer also proposed a self-attention distillation operation to dominate the attention scores in the stacked layers, further reducing the overall space complexity. The Informer model has been successfully applied to multiple fields, such as power prediction, drought prediction, medical devices, and air quality prediction, and has achieved remarkable results in these tasks.

[0006] Features such as Informer's dilated self-attention mechanism and self-attention distillation operation have significant application potential in the field of tunnel fire situation prediction, capable of providing faster and more accurate prediction results, which is of great significance for timely response and disaster prevention. However, the research and application of the Informer model in the field of tunnel fire prediction are still relatively limited and face the following problems: (1) The feature extraction ability of the sampling module of Informer is insufficient, which causes the model to be unable to deeply capture the relationships between tunnel fire time series data, thus affecting the fire situation prediction; (2) Informer runs relatively slowly as a whole when processing large-scale fire time series data, and the computational efficiency needs to be improved; (3) The high complexity and long-term dependence of tunnel fire time series data make it difficult for the model to capture the deep associations of the data, and the existing attention mechanisms fail to flexibly adapt to different data types, thus affecting the prediction accuracy. Summary of the Invention

[0007] The technical problem to be solved by the present invention: How to further improve the accuracy of tunnel fire prediction.

[0008] To solve the above technical problem, the present invention provides a tunnel fire prediction method based on the BiDS-Informer model, including:

[0009] Step 1, collect the fire time series data of the tunnel;

[0010] Step 2, construct a tunnel fire time series prediction model. The tunnel fire time series prediction model selects the BiDS-Informer model, and the BiDS-Informer model includes the encoder and decoder of the Informer model;

[0011] The input data of the encoder is the preprocessed tunnel fire time series data, and the features of the input data are extracted;

[0012] The decoder converts the output data of the encoder into a tunnel fire prediction result;

[0013] Downsampling modules are inserted between the multi-layer structures of the encoder and between the multi-layer structures of the decoder;

[0014] Step 3: Train the BiDS-Informer model and use an optimization algorithm to adjust the model parameters;

[0015] Step 4: Use the optimized BiDS-Informer model to predict the input original tunnel image data to obtain the final tunnel fire prediction result.

[0016] For the aforementioned tunnel fire prediction method based on the BiDS-Informer model, in Step 2, the encoder consists of multiple stacked attention layers, the attention layer includes a multi-head attention mechanism, and the downsampling module is located between the two attention layers;

[0017] The decoder consists of multiple attention layers, and the downsampling module is located between the two attention layers.

[0018] For the aforementioned tunnel fire prediction method based on the BiDS-Informer model, in Step 2, the downsampling module is a bidirectional downsampling module, including:

[0019] (1) The input data is sliced into a three-dimensional block structure of N×C×L, and local detail and long-range dependence features are captured through parallel 3×1 and 5×1 double-branch convolutions respectively; N, C, and L are the batch size, number of channels, and sequence length respectively;

[0020] (2) Then batch normalization processing and differential activation are performed respectively to enhance the non-linear expression; the formula for the batch normalization processing is:

[0021]

[0022] where and are the mean and standard deviation of the batch data respectively, and are the learnable parameter one and parameter two respectively, is a constant;

[0023] (3) Hierarchical downsampling is implemented using dynamic kernel size max pooling to form a multi-scale feature representation;

[0024] (4) The output features C1×L1 of channel 1 and the output features C2×L2 of channel 2 are reorganized into the standard format of the reconstructed channel N×C’×L’ through the channel dimension zero-padding alignment splicing method.

[0025] The aforementioned tunnel fire prediction method based on the BiDS-Informer model, in step 2, uses the Time Encoding Subtraction (TES) operation to optimize the dimensionality reduction of the encoder. The TES operation is expressed as:

[0026]

[0027] where, is the input feature, represents the time encoding function.

[0028] The aforementioned tunnel fire prediction method based on the BiDS-Informer model, in step 2, the decoder of the Informer model includes a multi-head diluted self-attention mechanism and a masked multi-head diluted self-attention mechanism. A linear layer is added to the front of the multi-head diluted self-attention mechanism and the masked multi-head diluted self-attention mechanism of the Informer model respectively to implement the dynamic attention mechanism.

[0029] The aforementioned tunnel fire prediction method based on the BiDS-Informer model, in the dynamic attention mechanism, first, the input Q, K, and V are projected to the multi-head dimension through linear projection. Each attention head corresponds to a separate subspace, which is represented by the following formula:

[0030]

[0031] where, 、 and are the projection matrices of the query, key, and value respectively, 、 and are the projections of the query, key, and value respectively. The projection operation converts the shape of the input data to or , where B is the batch size, L is the query length, S is the key and value length, and H is the number of attention heads.

[0032] The aforementioned tunnel fire prediction method based on the BiDS-Informer model selects more representative keys through the adaptive sampling network and calculates the sampling weights through the following formula:

[0033]

[0034]

[0035] where, and are the first and second parameters of the sampling network. By linearly transforming K, the and use function to normalize into adaptive sampling weights ;

[0036] Based on the adaptive sampling weights , through the selected keys, construct the sample set the top k most important simplified version sample sets in e, and then perform a dot product operation with Q. The formula is as follows:

[0037]

[0038]

[0039] where, is a list or array that contains the position information of the top k important keys in the original data, are the top k most important keys in the adaptive sampling weights, is the query after key screening, is the result obtained by operating the simplified query vector with the selected key vector;

[0040] Use the adaptive scaling factor to optimize the calculation of the attention score:

[0041]

[0042] where, represents the sigmoid function, and are the first and second parameters of the scaling network;

[0043] After the attention score is adjusted by the adaptive scaling factor, through function and function operations, the attention weight matrix A is obtained:

[0044]

[0045]

[0046] Initialize the context based on the masking mechanism. In the autoregressive task, the BiDS-Informer model cannot obtain future information;

[0047] Use the attention weight matrix A to weighted update V, accumulate the sequence information, and generate the final context output;

[0048] The multi-head outputs are merged back into the original feature dimension through a linear layer, making the output shape consistent with the input, and the updated context and attention weights are returned.

[0049] The aforementioned tunnel fire prediction method based on the BiDS-Informer model further includes a step of evaluating the optimized BiDS-Informer model after step 3. In the evaluation calculation of the BiDS-Informer model, it includes the mean absolute error MAE and the mean square error MSE, and the calculation formulas are as follows:

[0050]

[0051]

[0052] Among them, is the output of the prediction model, is the actual value, and n is the number of samples.

[0053] A computer system includes a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the above method.

[0054] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the above method.

[0055] The beneficial effects achieved by the present invention: The method of the present invention aims at the limitations of the existing downsampling module in processing multi-dimensional time series data, uses a bidirectional downsampling module to enhance the feature extraction ability, effectively improves the model's deep correlation capture of complex data, and then enhances the overall processing ability of the model, and processes sequence data more flexibly, quickly, and accurately, and then more accurately predicts various indicators of tunnel fires.

[0056] To improve the computational efficiency and running speed of the model, the present invention performs a dimensionality reduction operation on the time encoding in the encoder structure, so that it has higher speed and performance when processing large-scale data.

[0057] Aiming at the problem that the attention mechanism of the existing model cannot flexibly adapt to different data types, the present invention introduces a dynamic attention mechanism to realize the dynamic recognition and processing of multi-type data, thereby improving the prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 It is a module diagram of one layer of the attention layer in the encoder of the existing Informer;

[0059] Figure 2 It is a schematic diagram of the encoder architecture of the existing Informer model;

[0060] Figure 3 Schematic diagram of the bidirectional downsampling module in Embodiment 1 of the present invention;

[0061] Figure 4 Principle diagram of TES operation in Embodiment 2 of the present invention;

[0062] Figure 5 Schematic diagram of the structure of the Informer network;

[0063] Figure 6 Flowchart of the dynamic attention mechanism in Embodiment 3 of the present invention;

[0064] Figure 7 Schematic diagram of the FDS tunnel fire model;

[0065] Figure 8 Process flowchart of a random case after the database reconstruction and shuffling operation in Embodiment 3 of the present invention;

[0066] Figure 9 Bar chart of losses at different prediction lengths on the tunnel dataset in Embodiment 3 of the present invention;

[0067] Figure 10 Schematic diagram of the training loss at different prediction lengths in Embodiment 3 of the present invention. Detailed implementation manners

[0068] The technical solution of the present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0069] Embodiment 1

[0070] A tunnel fire prediction method based on the BiDS-Informer model, comprising:

[0071] Step 1, collecting the fire time series data of the tunnel, including temperature, carbon monoxide concentration, wind speed, etc., and preprocessing the time series data so that the input data is suitable for model training;

[0072] Step 2, constructing a tunnel fire time series prediction model. The tunnel fire time series prediction model selects the BiDS-Informer model (Bidirectional Down-sampling Informer model), and the BiDS-Informer model includes the encoder and decoder of the Informer model;

[0073] The input data of the encoder is the preprocessed tunnel fire time series data. The dilated attention mechanism is adopted to reduce the computational complexity of long sequence processing, and the features of the input data are extracted through multi-layer self-attention and feed-forward neural networks;

[0074] The decoder converts the output data of the encoder into the final tunnel fire prediction result. In the decoder, the prediction result is generated through multi-head self-attention, masked self-attention, and a feed-forward network, and a generative inference process is adopted to improve the inference speed;

[0075] Downsampling modules are inserted between the multi-layer structures of the encoder and between the multi-layer structures of the decoder;

[0076] Step 3: Train the BiDS-Informer model, and use an optimization algorithm to adjust the model parameters to reduce the prediction error. The optimization algorithm includes gradient descent;

[0077] Step 4: Evaluate the optimized BiDS-Informer model to ensure that its performance meets the requirements;

[0078] Step 5: Use the optimized BiDS-Informer model to predict the input original tunnel image data to obtain the final tunnel fire prediction result.

[0079] In step 2, the BiDS-Informer model includes the encoder and decoder of the existing technology Informer model;

[0080] In the existing technology Informer, the encoder and decoder respectively process the image data of the input time series and generate the prediction result. Downsampling modules are inserted between the multi-layer structures of the encoder and between the multi-layer structures of the decoder.

[0081] As Figure 1 shown, the encoder consists of multiple stacked attention layers. The attention layer includes a multi-head attention mechanism. The downsampling module is located between the two attention layers. After the output of each attention layer, a downsampling operation is used to reduce the sequence length;

[0082] The decoder also consists of multiple attention layers. The downsampling module is located between the two attention layers to reduce the sequence length that needs to be processed during the decoding process.

[0083] As Figure 2 shown, the Informer model integrates a convolutional layer and a max-pooling layer between the self-attention layers, aiming to reduce the length of the input sequence. The convolutional layer is configured with a kernel size of 3 and a stride of 1 to enhance the ability to capture local context features. The mathematical expression is:

[0084]

[0085] Among them, is the weight of the convolutional kernel, is the bias term, is a convolution operation, is an activation function, is a one-dimensional convolution.

[0086] Subsequently, the max pooling layer (kernel size 3, stride 2) further refines the context features, highlighting the key local information and providing a more accurate feature map for the subsequent self-attention layer. The operation of the max pooling layer can be expressed as:

[0087]

[0088] where, represents the elements in the sequence from position to in the sequence, is one-dimensional max pooling.

[0089] Traditional convolutional layers have limitations in the field of time series prediction: the receptive field only grows linearly with the network depth, which limits the ability to process ultra-long sequences and may lead to information leakage, causing ineffective calculations. In addition, traditional convolutional layers do not fully consider the temporal dependence of time series, which may lead to the leakage of future information during prediction, thus affecting the performance of the Informer model in processing extremely long sequences.

[0090] To address the problems encountered by the Informer model in processing long-sequence time series prediction tasks, including limited processing capacity, single feature extraction method, and insufficient robustness, the downsampling module of the model was innovatively improved by increasing the number of channels in the downsampling module to enhance the parallel processing ability of the model, enabling the model to process large amounts of data more efficiently, thereby improving the overall computational performance.

[0091] In step 2, the downsampling module is a bidirectional downsampling module, such as the improved bidirectional downsampling module shown in Figure 3 , which realizes efficient feature extraction through the collaborative mechanism of block division, multi-scale convolution, and cross-branch feature fusion:

[0092] (1) Block division: The input data is sliced into a three-dimensional block structure of N×C×L, and local detail and long-range dependence features are captured respectively through parallel 3×1 and 5×1 double-branch convolutions; the stride of the double-branch convolution is 2; N, C, and L are the batch size, number of channels, and sequence length respectively;

[0093] (2) Then, batch normalization and differential activation (ReLU / LeakyReLU) are performed respectively to enhance the non-linear expression; the formula for the batch normalization is:

[0094]

[0095] where, and are the mean and standard deviation of the batch data respectively, and are learnable parameter one and parameter two, is a small constant added to prevent zero error; Batch normalization helps to accelerate the training process and improve the stability of the model.

[0096] (3) Multi-scale convolution: Hierarchical downsampling is implemented using dynamic kernel size max pooling (2×1 and 3×1) to form multi-scale feature representations;

[0097] (4) Cross-branch feature fusion. The output features C1×L1 of channel 1 and the output features C2×L2 of channel 2 are reorganized into the standard format of the reconstructed channel N×C’×L’ through the method of zero-padding alignment and splicing in the channel dimension.

[0098] The improved bidirectional downsampling module reduces the computational load by blocks, strengthens feature diversity through multi-scale convolution, condenses key information through differential pooling, and maintains feature integrity through intelligent splicing. Finally, it maximally retains spatial-channel information during the dimensionality reduction process, providing a high-density, multi-scale fused feature representation for downstream tasks. Different sizes of convolutional kernels and various activation functions are used, enabling the model to capture the features of the data from multiple perspectives. The convolution operation allows the model to extract features at different scales of the time series, while the activation function introduces non-linearity, enhancing the expressive power of the model. Through parallel processing and feature fusion strategies, the limitations of the traditional Informer downsampling module are effectively overcome. The feature fusion strategy allows the model to integrate information at different levels, thus providing more accurate prediction results in the long-sequence time series prediction task, not only enhancing the robustness of the Informer model but also improving its efficiency in processing extremely long sequences.

[0099] Example 2

[0100] A tunnel fire prediction method based on the BiDS-Informer model, comprising:

[0101] Step 1, collect the fire time series data of the tunnel, including temperature, carbon monoxide concentration, wind speed, etc., and preprocess the time series data to make the input data suitable for model training;

[0102] Step 2, construct a tunnel fire time series prediction model. The tunnel fire time series prediction model selects the BiDS-Informer model, and the BiDS-Informer model includes the encoder and decoder of the Informer model;

[0103] The input data of the encoder is the preprocessed tunnel fire time series data. The dilated attention mechanism is adopted to reduce the computational complexity of long sequence processing, and the features of the input data are extracted through multi-layer self-attention and feed-forward neural networks.

[0104] The decoder converts the output data of the encoder into the final tunnel fire prediction result. In the decoder, the prediction result is generated through multi-head self-attention, masked self-attention and feed-forward network, and a generative inference process is adopted to improve the inference speed.

[0105] Downsampling modules are inserted between the multi-layer structures of the encoder and between the multi-layer structures of the decoder.

[0106] Step 3: Train the BiDS-Informer model, and use the optimization algorithm to adjust the model parameters to reduce the prediction error. The optimization algorithm includes gradient descent.

[0107] Step 4: Evaluate the optimized BiDS-Informer model to ensure that its performance meets the requirements.

[0108] Step 5: Use the optimized BiDS-Informer model to predict the input original tunnel image data to obtain the final tunnel fire prediction result.

[0109] On the basis of Embodiment 1, in Step 2, the encoder is optimized for dimensionality reduction by using the Time Encoding Subtraction (TES) operation.

[0110] The Informer model is an efficient architecture for long sequence time series data. Although it performs excellently in capturing long-distance dependencies, its training and inference time is relatively long when dealing with large-scale datasets, which limits the feasibility of real-time applications.

[0111] In the Time Encoding Subtraction (TES), the complexity of the model is reduced by preprocessing the input data to subtract the time encoding. The data dimension of the self-attention layer is reduced, thereby reducing the computational complexity and significantly improving the training and inference speed of the Informer model.

[0112] The TES operation can be expressed as:

[0113]

[0114] where is the input feature, represents the time encoding function.

[0115] The TES method reduces the number of matrix multiplication operations for each attention head, decreases memory occupancy, and improves running efficiency. In addition, TES enhances the resource utilization rate of the model, reduces the demand for hardware resources such as GPUs, allows for more parallel computing, and further boosts the training speed. Meanwhile, the TES operation does not affect the prediction accuracy of the model, achieving high prediction performance while reducing the computational amount.

[0116] In a tunnel fire scenario, since a fire is an unexpected event and has no direct association with traditional global time encoding. By reducing the dimension to remove these global time factors, the model can focus more on local time patterns related to the fire, improving adaptability and prediction accuracy. This optimization reduces irrelevant features, decreases model complexity, enhances computational efficiency, enables the model to capture critical time periods more accurately, avoids overfitting, and predicts unexpected events more effectively.

[0117] Embodiment 3

[0118] A tunnel fire prediction method based on the BiDS-Informer model, comprising:

[0119] Step 1, collect the fire time series data of the tunnel, including temperature, carbon monoxide concentration, wind speed, etc., and preprocess the time series data so that the input data is suitable for model training;

[0120] Step 2, construct a tunnel fire time series prediction model. The tunnel fire time series prediction model selects the BiDS-Informer model, and the BiDS-Informer model includes the encoder and decoder of the Informer model;

[0121] The input data of the encoder is the preprocessed tunnel fire time series data. The dilated attention mechanism is adopted to reduce the computational complexity of long sequence processing, and the features of the input data are extracted through multi-layer self-attention and feed-forward neural networks;

[0122] The decoder converts the output data of the encoder into the final tunnel fire prediction result. In the decoder, the prediction result is generated through multi-head self-attention, masked self-attention, and feed-forward networks, and a generative inference process is adopted to improve the inference speed;

[0123] Insert downsampling modules between the multi-layer structures of the encoder and between the multi-layer structures of the decoder;

[0124] Step 3, train the BiDS-Informer model, and use an optimization algorithm to adjust the model parameters to reduce the prediction error. The optimization algorithm includes gradient descent;

[0125] Step 4, evaluate the optimized BiDS-Informer model to ensure that its performance meets the requirements;

[0126] Step 5: Use the optimized BiDS-Informer model to predict the input original in-tunnel image data to obtain the final tunnel fire prediction result.

[0127] The self-attention mechanism performs excellently in deep learning, especially when dealing with long sequence data. However, traditional methods have some deficiencies. For example, they uniformly process all input data and fail to fully consider the importance of the data, resulting in low processing efficiency of redundant information. In addition, the fixed attention range limits the adaptive ability of the model. Therefore, a dynamic self-attention mechanism is proposed. By introducing a content-based adaptive sampling network and an adaptive scaling factor, the model can dynamically adjust the attention strategy and weight distribution according to the input features. This improves the computational efficiency and enhances the adaptability and expressive ability of the model in different tasks, better coping with the challenges of complex sequence data processing.

[0128] As Figure 5 shown, based on Embodiment 2, in Step 2, the decoder of the Informer model includes a multi-head diluted self-attention mechanism and a masked multi-head diluted self-attention mechanism. A linear layer is added to the front section of the multi-head diluted self-attention mechanism and the masked multi-head diluted self-attention mechanism of the Informer model respectively to implement the dynamic attention mechanism.

[0129] In the dynamic attention mechanism, first, the input queries (Q), keys (K), and values (V) are transformed to the multi-head dimension through linear projection. Each attention head corresponds to a separate subspace, which is represented by the following formula:

[0130]

[0131] where , and are the projection matrices of the query, key, and value respectively. , and are the projections of the query, key, and value respectively. The projection operation transforms the shape of the input data into or , where B is the batch size, L is the query length, S is the key and value length, and H is the number of attention heads.

[0132] Different from the random sampling method of the traditional attention mechanism, in this embodiment, a more representative key is selected through the adaptive sampling network, and the sampling weight is calculated by the following formula:

[0133]

[0134]

[0135] Among them, and are the first and second parameters of the sampling network. By performing a linear transformation on K, is obtained for generating the sampling weights, and is normalized by the function to the adaptive sampling weights

[0136] . This improvement enables the model to adaptively adjust the sampling weights according to the input data. In the traditional self-attention mechanism, this step is usually achieved by randomly selecting sampling points in K, but this method lacks representativeness. The improved mechanism selects points through learning, which is more accurate and effective. Based on the adaptive sampling weights , the sample set is constructed by selecting the first k most important simplified version sample sets in e, and then a dot product operation is performed with Q. The formula is as follows:

[0137]

[0138]

[0139] Among them, is a list or array that contains the position information of the first k important keys in the original data, are the first k most important keys in the adaptive sampling weights, is the query after being filtered by the key, is the result obtained by operating the simplified query vector with the selected key vector. Compared with the model without using the dynamic attention mechanism, this improved version ensures the representativeness of the sampling points and avoids the problem of unimportant keys caused by randomness.

[0140] Optimize the calculation of the attention score using the adaptive scaling factor :

[0141]

[0142] Among them, represents the sigmoid function, and are the first and second parameters of the scaling network. This scaling factor is dynamically adjusted according to the mean of Q, enabling the model to adaptively scale the attention score when processing different input data. Compared with using a fixed scaling factor in the attention mechanism, the improved adaptive scaling factor can more flexibly handle different data distributions, improving the robustness and adaptability of the model.

[0143] After the improved attention scores are adjusted by the adaptive scaling factor, through function and function operations, the attention weight matrix A is obtained:

[0144]

[0145]

[0146] Subsequently, the context is initialized based on the masking mechanism. In the autoregressive task, the BiDS-Informer model cannot obtain future information. Then, V is weighted and updated using A to accumulate sequence information and generate the final context output. Finally, the multi-head output is merged back to the original feature dimension through a linear layer to make the output shape consistent with the input, and the updated context and attention weights are returned.

[0147] As Figure 6 shown in the flowchart of the dynamic attention mechanism. Compared with the unimproved version, this dynamic attention mechanism significantly improves the accuracy of the output and the adaptability to different input data through an adaptive method. The innovation in the sampling and scaling process of the improved mechanism makes the model more intelligent and efficient, thereby reducing the computational complexity while maintaining performance, and showing higher robustness and flexibility in practical applications.

[0148] The method of the present invention is described below through experiments. The experimental environment configuration is shown in Table 1:

[0149] Table 1 Experimental Environment Configuration

[0150] Name Configuration Information Operating System Microsoft Windows 11 Home Chinese Edition Development Language Python 3.12.5 Framework Pytorch 2.3.1 + cuda 11.8 CPU 12th Gen Intel(R) Core(TM) i5-12450H GPU NVIDIA GeForce RTX 4050 Laptop GPU

[0151] Dataset and preprocessing: To evaluate the performance improvement of the BiDS-Informer model, three ETT datasets are selected for comprehensive comparative experiments. The following is the description of the ETT datasets. The ETT datasets, namely the power transformer temperature, are specifically used for the performance test of long-sequence prediction models. The datasets include various attribute data of the transformer, including seven variables such as load, temperature, and humidity, and are divided according to the sampling interval and data characteristics. Among them, ETTh1 is sampled once per hour, and ETTm1 and ETTm2 are sampled once every 15 minutes, and the overall time span of the ETT datasets is one and a half years.

[0152] The ECL dataset, namely the electricity load, is selected to evaluate the performance of the long-sequence time series prediction model. This dataset records two years of hourly electricity consumption information and has 321 variables, which is a typical dataset for load forecasting and multivariate time series analysis.

[0153] The dataset used is the tunnel fire simulation dataset, which is generated by the FDS software. A specific scenario is selected, that is, a certain section of the tunnel loses power after a fire breaks out. A full-scale tunnel model based on a shielded tunnel is constructed by FDS. The tunnel is 180 m long and 2.75 m in radius. The fire source is modeled as a cube with dimensions of 2m×2m×2m and is defined as "BURNER" with an area of 4m². In addition, fire sources are set sequentially along the path. At the same time, sensors such as thermocouples, CO sensors, and anemometers are placed regularly in sequence to monitor the environmental conditions. To simplify the calculation, the tunnel wall uses the default concrete material with a thickness of 0.3 m. At the entrance end of the tunnel, when there is a ventilation wind speed, the boundary condition is set to "SUPPLY". Otherwise, it is set to "OPEN", that is, the exit end. In addition, the environmental temperature is 20°C, and the simulation time is 360 s, meeting the tunnel design specification GB 50157 - 2013.

[0154] After generating the tunnel fire time series database through FDS numerical simulation, as Figure 8 shown is the process flow chart of a random case after reconstructing and shuffling the database. The dataset of 100 scenarios is randomly divided into a training set, a validation set, and a test set, with proportions of 60%, 20%, and 20% respectively.

[0155] Table 2 shows the statistical parameters of the tunnel fire numerical simulation database after division, including sensor data such as temperature, carbon monoxide concentration, and wind speed. From the "data length" column, it can be seen that the length of this database is 2 to 8 times that of the public dataset ETT, indicating that the data volume is sufficient, which helps to improve the training effect and prediction accuracy of the model. The "maximum value" and "minimum value" columns provide the maximum and minimum values of the corresponding dataset, which are crucial for understanding the data range and the impact of extreme values. The "average value" and "standard deviation" columns show that the average values of each subset are relatively stable and the standard deviation is moderate, indicating that the data shows a reasonable range of variation, which helps the model to capture potential trends without being disturbed by excessive noise. In addition, the "skewness" value in the table indicates that the data has a certain skewness, which affects the prediction ability of the model. Skewness means that the data distribution is asymmetric, resulting in prediction deviations of the model in some cases. Therefore, in subsequent analysis, considering appropriate transformation of the data to reduce skewness will help to improve the robustness and accuracy of the model.

[0156] In summary, the selected dataset has a rich sample size, reasonable statistical characteristics, and adjustable skewness. These characteristics together support the subsequent prediction tasks, ensuring that the model can effectively identify and predict the risks of tunnel fires.

[0157] Table 2 shows the statistical characteristics of the dataset

[0158]

[0159] In model evaluation and error calculation, including the mean absolute error (MAE) and the mean squared error (MSE), the lower the metric value, the smaller the error of the model. The metric calculation formulas are as follows:

[0160]

[0161]

[0162] Among them, is the output of the prediction model, is the actual value, n is the number of samples. The closer the values of MAE and MSE are to 0, the more accurate the prediction.

[0163] To verify the performance improvement of the BiDS-Informer model, six time series prediction methods were selected for multivariate comparison experiments, including the typical self-attention variant Informer, the classical attention Transformer, FNet that uses Fourier transform to improve computational efficiency, and the most classical long sequence time series prediction model LSTM, which can comprehensively evaluate the performance of different models in multivariate time series prediction.

[0164] The model of the present invention is trained using the L2 loss function, the optimizer uses ADAM, and the initial learning rate is set to 10-4. The batch size is 128, the number of training iterations is 15, and an early stopping strategy is introduced. All experiments are implemented under the PyTorch framework and run on an NVIDIA GeForce RTX 14GB GPU, and each experiment is repeated three times. The model structure contains 3 encoder layers and 1 decoder layer to balance performance and computational efficiency. In terms of hyperparameter settings, the prediction window size L is set to 12, 24, and 48, and the input sequence length and the start token length of Informer are set to 96 and 48 respectively. The sparsity parameter u controls the computational amount of the sparse self-attention mechanism and determines the number of keys selected each time. When u is small, the computational efficiency is higher, but some accuracy may be sacrificed; when u is large, more information can be captured, but the computational amount will also increase. Therefore, u can be adjusted according to the task requirements to achieve a balance between efficiency and accuracy.

[0165] Table 3 summarizes the multivariate prediction results of five methods on five datasets. On the public datasets, the performance of BiDS-Informer at prediction lengths of 24 seconds and 48 seconds is first presented. To further evaluate the performance of BiDS-Informer on the tunnel fire dataset, the performance metrics for a prediction length of 12 seconds are added. The best results predicted by the model are highlighted in bold in the table, and the second-best results are underlined. The analysis results show that BiDS-Informer exhibits the optimal performance on all datasets, followed by the classic Informer and LSTM models.

[0166] On the three ETT datasets, compared with Informer, the average MSE performance of BiDS-Informer is improved by 4.43% (0.631→0.603), 6.82% (0.425→0.356), and 15.28% (0.216→0.183) respectively, and its average MAE performance is improved by 4.97% (0.604→0.574), 4.77% (0.419→0.399), and 9.31% (0.333→0.302) respectively. On the ECL dataset, compared with the second-best performing LSTM, the average MSE and MAE performance of BiDS-Informer are improved by 3.7% (0.324→0.312) and 1.77% (0.396→0.389) respectively. On the tunnel fire dataset, the prediction effect of BiDS-Informer is also the most significant, with the average MSE and MAE performance improved by 17.45% (0.361→0.298) and 26.3% (0.361→0.311) respectively compared with the second-best performing Informer at the three prediction lengths. The BiDS-Informer model demonstrates good performance in multivariate time series prediction, which benefits from its advantages in prediction accuracy, stability, and adaptability. Whether on public datasets or datasets in specific fields such as tunnel fires, the model of the present invention can provide high-quality prediction results. The algorithm optimization of BiDS-Informer enables it to be efficient in capturing the complex patterns of time series data and maintain the consistency of prediction results in different environments, which is crucial for application scenarios that rely on prediction results in the long term. In addition, the low-error prediction ability of the model further proves its superiority in revealing the internal laws of time series data, ensuring its reliability in practical applications.

[0167] Table 3 shows the prediction results of multivariate long sequence time series on five datasets

[0168]

[0169] To verify the performance of different time series prediction models on the tunnel fire dataset, the last few columns of data in Table 3 above are shown below Figure 9 as follows Figure 9 The following shows the error index results of 5 time series prediction models at different prediction lengths of 12s, 24s, and 48s. In the figure, (a) and (b) are the comparisons of the MSE and MAE indicators respectively. The abscissa is different models, and the ordinate is the evaluation index. It can be seen from the figure that as the prediction length increases, the corresponding prediction error index increases significantly. BiDS-Informer shows the lowest MAE and MSE at all prediction lengths (12s, 24s, 48s), which are 0.237 and 0.483 (12s), 0.299 and 0.715 (24s), 0.358 and 1.068 (48s) respectively, showing its superiority in short-term and medium-term predictions. In contrast, the Transformer and Fnet models perform relatively close at each prediction length, but the overall error is slightly higher. Especially at the 48s prediction length, the MAE and MSE are as high as 0.368 and 0.695, indicating their deficiencies in long-term predictions. The LSTM model performs the worst, especially the MSE at the 48s prediction length is as high as 1.068, showing its limitations in dealing with complex time series. Generally speaking, as the prediction length increases, the overall error generally rises, especially for the Informer and LSTM models. The error at 48s is significantly higher than that at 12s and 24s. Although the error of BiDS-Informer increases at 48s, compared with other models, it still maintains a relatively low error level, indicating its certain robustness in long-term predictions. In summary, the BiDS-Informer model performs excellently in the tunnel fire prediction task, providing an important reference for future research and applications.

[0170] When training a model, MSE is usually used as the loss function to promote the rapid convergence of the model. As follows Figure 10 The following shows the changes in the mean squared error MSE loss with the number of training epochs when different models predict the tunnel dataset. Figures (a), (b), and (c) show the MSE loss at prediction lengths of 12s, 24s, and 48s respectively.

[0171] As can be seen from Figure (a), as the number of training epochs increases, the training losses of all models show an obvious downward trend. The BiDS-Informer model (red line) rapidly reduces the loss at the beginning of training and stabilizes at a level close to 0.1, showing its superior convergence and learning ability. In contrast, the loss of the LSTM (blue line) model decreases slowly and finally stays at a relatively high loss level, indicating that its learning effect in this task is not as good as other models.

[0172] In Figure (b), the overall trend is similar to that in Figure (a), but the loss values of each model are generally higher. Nevertheless, BiDS-Informer still maintains a relatively low loss, demonstrating its robustness under longer prediction lengths. The losses of the Informer and Transformer models also decrease significantly in the initial stage of training, but the convergence speed slows down when approaching 0.5, showing certain learning limitations.

[0173] In Figure (c), the changing trend of the loss curve is consistent with the previous two Figure 1 ones, but the overall loss level further increases. In particular, the loss values of the Fnet and LSTM models are significantly higher than those of other models. Although the loss of BiDS-Informer is still relatively low at a prediction length of 48s, its convergence speed slows down slightly, indicating that it may face greater challenges in long-term predictions.

[0174] In summary, BiDS-Informer demonstrates superior training performance at different prediction lengths, with its loss value continuously lower than that of other models, proving its effectiveness and reliability in tunnel fire prediction tasks.

[0175] Compare the performance of the dynamic attention mechanism, encoder dimensionality reduction optimization, and bidirectional downsampling module on the tunnel fire dataset. Table 4 shows the results of the ablation experiment, which is used to verify the impact of different modules on the prediction results of tunnel fire data. First, the impact of encoder dimensionality reduction optimization on the prediction results is relatively small. When predicting 12 time steps, the error of the model is the smallest, with the MAE and MSE being 0.2386 and 0.2729 respectively, while the error of the final model is 0.2375 and 0.2994. In addition, removing the dynamic attention mechanism has a significant negative impact on the results. When predicting the temperature, wind speed, and carbon monoxide concentration in the next 48 seconds, the errors of each metric are 0.0623, 0.0544, and 0.0500 higher than those of the original model respectively. In summary, the impacts of different modules on the final prediction results vary: the impact of encoder dimensionality reduction optimization is relatively small, the impact of the dynamic attention mechanism is relatively significant, and the impact of the bidirectional downsampling module is between the two.

[0176] Table 4 Results of the ablation experiment

[0177]

[0178] The proposed BiDS-Informer in the present invention is a multivariate time series prediction model based on Informer. It uses a dynamic attention mechanism to adaptively adjust the attention allocation, thereby capturing more critical features, enabling BiDS-Informer to flexibly handle the diversity of input data and improve the accuracy of output data. In addition, we also designed a new bidirectional downsampling module, which can efficiently extract information at different time scales, significantly enhancing the timeliness of the model. At the same time, the encoder dimension reduction improves the expression ability of long sequence data while reducing the computational cost. To verify the generalization ability of this method, we conducted extensive evaluations on three ETT power datasets and the ECL dataset respectively. On different prediction lengths of the three ETT datasets, the average MSE and MAE of BiDS-Informer are 0.429 and 0.425 respectively. At the same time, on the ECL dataset, the average MSE and MAE under different prediction lengths are 0.327 and 0.399 respectively. These outstanding results indicate that the proposed BiDS-Informer can effectively mine different types of time series data to achieve effective prediction.

[0179] A computer system includes a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the method as described above.

[0180] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, it implements the steps of the method as described above.

[0181] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0182] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one or more flows and / or blocks Figure 1The functions specified in one or more boxes.

[0183] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the steps of the functions specified in one Figure 1 one process or more processes and / or boxes Figure 1 step of the functions specified in one or more boxes.

[0184] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A tunnel fire prediction method based on BiDS-Informer model, characterized in that: include: Step 1, collecting tunnel fire time series data; Step 2, construct a tunnel fire time series prediction model, the tunnel fire time series prediction model uses the BiDS-Informer model, and the BiDS-Informer model includes an encoder and a decoder of the Informer model; The encoder input data is pre-processed tunnel fire time series data, and the features of the input data are extracted; The decoder converts the output data of the encoder into a tunnel fire prediction result; A downsampling module is inserted between the multi-layer structure of the encoder and between the multi-layer structure of the decoder; Step 3, train the BiDS-Informer model and adjust the model parameters using the optimization algorithm; Step 4: Use the optimized BiDS-Informer model to predict the input original tunnel image data to obtain the final tunnel fire prediction result.

2. A tunnel fire prediction method based on BiDS-Informer model according to claim 1, characterized in that: In step 2, the encoder is composed of a plurality of stacked attention layers, the attention layer includes a multi-head attention mechanism, and the downsampling module is located between the two attention layers; The decoder consists of multiple attention layers, with the downsampling module located between the two attention layers.

3. A tunnel fire prediction method based on BiDS-Informer model according to claim 1 or 2, characterized in that: In step 2, the downsampling module is a bidirectional downsampling module, comprising: (1) The input data is divided into a three-dimensional block structure of N×C×L, and the local details and long-range dependency features are captured by parallel 3×1 and 5×1 dual-branch convolutions respectively; N, C, and L are the batch size, number of channels, and sequence length respectively; (2) Then batch normalization and differential activation are performed to enhance nonlinear expression. The formula for batch normalization is: ; in, and are the mean and standard deviation of the batch data, and They are learnable parameters one and two, is a constant; (3) Using dynamic kernel size maximum pooling to implement layered downsampling to form multi-scale feature expression; (4) The output features C1×L1 of channel 1 and the output features C2×L2 of channel 2 are reorganized into the standard format of the reconstructed channel N×C'×L' through the channel dimension zero-padding alignment splicing method.

4. The tunnel fire prediction method based on the BiDS-Informer model according to claim 1 is characterized in that: In step 2, the encoder is optimized by using the temporal encoding subtraction TES operation, which is expressed as: ; in, are input features, Represents a time encoding function.

5. The tunnel fire prediction method based on the BiDS-Informer model according to claim 1 is characterized in that: In step 2, the decoder of the Informer model includes a multi-head diluted self-attention mechanism and a masked multi-head diluted self-attention mechanism. Linear layers are added to the front end of the multi-head diluted self-attention mechanism and the masked multi-head diluted self-attention mechanism of the Informer model to realize the dynamic attention mechanism.

6. A tunnel fire prediction method based on BiDS-Informer model according to claim 5, characterized in that: In the dynamic attention mechanism, first, the input Q, K, and V are transformed into multi-head dimensions through linear projection, and each attention head corresponds to a separate subspace, which is expressed by the following formula: ; in, , and are the projection matrices for query, key, and value, respectively, , and They are query, key, and value projections. The projection operation transforms the input data shape into or , where B is the batch size, L is the query length, S is the length of the key and value, and H is the number of attention heads.

7. A tunnel fire prediction method based on BiDS-Informer model according to claim 6, characterized in that: More representative keys are selected through an adaptive sampling network, and the sampling weight is calculated by the following formula: ; ; in, and are the parameters 1 and 2 of the sampling network. By linearly transforming K, we can obtain the parameters used to generate the sampling weights. and use Function normalization to adaptive sampling weights ; Based on adaptive sampling weight , by selecting Key, construct a sample set The first k most important simplified version samples in e , and then perform a dot product operation with Q, the formula is as follows: ; ; in, is a list or array containing the location information of the first k important keys in the original data. are the top k most important keys in the adaptive sampling weights, It is through Query after key filtering, It is the result obtained by operating the simplified query vector with the selected key vector; Using Adaptive Scaling Factor Optimize the calculation of attention scores: ; in, represents the sigmoid function, and are parameters one and two of the scaling network; The attention scores are adjusted by the adaptive scaling factor. Functions and Function operation to obtain the attention weight matrix A: ; ; Based on the mask mechanism to initialize the context, the BiDS-Informer model cannot obtain future information in the autoregressive task; Use the attention weight matrix A to perform weighted updates on V, accumulate sequence information, and generate the final context output; The multi-head outputs are merged back to the original feature dimension through a linear layer to make the output shape consistent with the input, and the updated context and attention weights are returned.

8. The tunnel fire prediction method based on BiDS-Informer model according to claim 1 is characterized in that: After step 3, the step of evaluating the optimized BiDS-Informer model is also included. In the evaluation calculation of the BiDS-Informer model, the mean absolute error MAE and the mean square error MSE are included. The calculation formula is as follows: ; ; in, is the output of the prediction model, is the actual value, and n is the number of samples.

9. A computer system, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.