A tunnel fire inversion method based on multi-modal fusion and physical constraint

By employing a multimodal fusion and physical constraint-based tunnel fire inversion method, and combining temperature sensor data and image data, a multimodal inversion network model is established. This solves the real-time and accuracy problems of tunnel fire source inversion, achieves accurate inversion of fire source power and location, and demonstrates stability and generalization in scenarios with missing multi-source data.

CN121351635BActive Publication Date: 2026-03-03CHINA UNIV OF MINING & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511897291.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-03-03
Estimated Expiration
2045-12-16

AI Technical Summary

Technical Problem

Existing methods for inverting tunnel fire sources involve large computational loads and long response times. Furthermore, methods based on deep learning models lack consistent modeling of fire evolution timelines and ignore physical constraints, resulting in unreasonable prediction results and poor generalization.

Method used

A tunnel fire inversion method based on multimodal fusion and physical constraints is adopted. By collecting time-series data from distributed temperature sensors, temperature distribution images, and physical parameters, a multimodal training dataset is constructed. A multimodal inversion network model is established, which includes time feature extraction, image encoding, physical parameter embedding, and gating fusion. A physical constraint loss and anti-information leakage mechanism are designed. The network model is trained using the multimodal training dataset to invert the power and location of the fire source in real time.

Benefits of technology

It achieves accurate real-time inversion of fire source power and location, conforms to thermal laws, improves inversion accuracy and physical consistency, and adapts to the stability and generalization performance under the condition of missing multi-source data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121351635B_ABST
    Figure CN121351635B_ABST
Patent Text Reader

Abstract

The application provides a tunnel fire inversion method based on multi-modal fusion and physical constraints, aiming at solving the problem that the fire source position and power are difficult to be accurately inverted in real time in traditional fire monitoring. The method collects distributed temperature sensor data, temperature distribution images and physical parameters in the tunnel, constructs a space-time feature fusion network, designs an information leakage prevention training mechanism, adopts a mode random discarding and sensor shielding technology to avoid overfitting of the model to a specific input mode, introduces a physical constraint loss function, combines a temperature gradient, a maximum temperature rise and a longitudinal attenuation law to ensure the physical rationality of the inversion result, and adopts a coarse-precision two-stage position prediction architecture to realize meter-level precision positioning of the fire source position. The application can invert the fire source power and position in real time and accurately, and provides reliable technical support for intelligent monitoring and emergency rescue of tunnel fires.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of disaster prevention and mitigation technology in civil engineering, and in particular to a method for inverting tunnel fire conditions based on multimodal fusion and physical constraints. Background Technology

[0002] Tunnel fires are one of the most serious types of disasters in underground transportation engineering, characterized by their suddenness, rapid spread, high heat flux density, and highly toxic smoke. Once a fire occurs, the high-temperature smoke inside the tunnel will rapidly spread longitudinally within a short period of time and form a complex backflow zone, making it difficult to determine the location of the fire source and resulting in a highly non-uniform temperature field distribution, which seriously affects firefighting and personnel evacuation.

[0003] Current methods for deriving tunnel fire sources mainly fall into two categories: one is numerical simulation methods based on computational fluid dynamics (CFD), such as Fire Dynamics Simulator (FDS). These methods can obtain high-precision temperature field distributions by refining the mesh and setting boundary conditions, but they involve large computational loads and long response times, making them unsuitable for emergency scenarios. The other category is methods based on empirical models or sensor measurement points. These methods typically rely on limited temperature or smoke sensor data, estimating the fire source location and power through formula fitting or multiple iterations. However, their results are easily affected by noise, boundary conditions, and sensor placement, resulting in poor stability.

[0004] With the development of artificial intelligence technology, the application of deep learning models in fire inversion is gradually increasing. Some studies use convolutional neural networks (CNNs) to extract spatial features from temperature field images or use recurrent neural networks (LSTMs) to analyze sensor time-series data to predict fire source parameters. However, these models generally have three shortcomings: First, they lack temporal consistency modeling of fire evolution, making it difficult to reflect the dynamics of fire source development; second, they ignore physical constraints such as heat conduction and energy conservation, leading to physically unreasonable prediction results; and third, the training process is prone to information leakage, meaning the model learns implicit patterns strongly correlated with labels rather than fire patterns, resulting in poor generalization under unseen conditions. Summary of the Invention

[0005] This invention aims to at least partially solve one of the technical problems in related technologies. Therefore, the first objective of this invention is to propose a tunnel fire inversion method based on multimodal fusion and physical constraints, capable of retrieving fire source power and location in real time and accurately, providing reliable technical support for intelligent monitoring and emergency rescue of tunnel fires.

[0006] To achieve the above objectives, a first aspect of the present invention proposes a tunnel fire inversion method based on multimodal fusion and physical constraints, the method comprising:

[0007] S1: Collect time-series data, temperature distribution images, and physical parameters from distributed temperature sensors inside the tunnel to construct a multimodal training dataset;

[0008] S2, establish a multimodal inversion network model that includes temporal feature extraction, image encoding, physical parameter embedding and gating fusion;

[0009] S3, Design a physical constraint loss and information leakage prevention mechanism, and train the multimodal inversion network model using a multimodal training dataset;

[0010] S4 utilizes the trained multimodal inversion network model to invert the power and location of the fire source in real time and outputs the fire inversion results.

[0011] According to an embodiment of the present invention, step S1 includes:

[0012] S11, temperature sensor data, temperature cloud map images and physical parameters under different fire source conditions in the tunnel are obtained through FDS simulation;

[0013] S12, Perform sliding window processing on the temperature sensor data to extract the current temperature and its first-order difference features, and form a time-series feature matrix;

[0014] S13, standardize and resize the temperature distribution image;

[0015] S14, Physical parameter normalization and leakage prevention measures.

[0016] According to an embodiment of the present invention, step S11 includes:

[0017] S111, construct a tunnel numerical model of a preset size, and determine the number and corresponding location of fire sources, the configuration range of fire source power, and the configuration parameters of longitudinal ventilation in the tunnel numerical model;

[0018] S112, determine the range of the fire source, the power growth of the fire source follows the t² fire model, the total simulation time is the first preset time; the fuel type is heptane, and the carbon monoxide generation rate and smoke particle production rate in the combustion products of the fuel are set respectively;

[0019] S113, a dynamic grid partitioning strategy is used to partition the tunnel space;

[0020] S114, the ventilation system includes two sets of jet fans, symmetrically arranged in the tunnel, each set of jet fans contains 2 fans; the volumetric flow rate of the air supply duct is precisely controlled by the HVAC system built into the FDS software to achieve the target wind speed condition.

[0021] S115, along the longitudinal centerline of the tunnel, multiple temperature sensors are deployed at preset intervals. Each sensor records temperature data, with a sampling interval of 1 second. At the same time, a temperature monitoring plane is set at the longitudinal center of the tunnel to export temperature cloud map data, with an image sampling interval of 1 second.

[0022] According to an embodiment of the present invention, step S12 includes:

[0023] S121, Set a sliding window to extract features from the time series data of each sensor using a sliding window of fixed length;

[0024] S122, for any time point t, extract the matrix of temperature readings from all sensors at the current time and the previous time, as shown in the following formula:

[0025] ;

[0026] Among them, T i (t) represents the temperature value at the i-th measuring point at time t, N s The number of sensors; the temperature feature matrix X at time t. temp (t) is a W-row N-row s A column matrix, expressed as follows:

[0027] ;

[0028] The first row of the matrix represents the global temperature at the previous time t-1, and the second row represents the global temperature at the current time t.

[0029] S123, calculate the temperature difference characteristics. The instantaneous rate of temperature change contains important information about the dynamic evolution of the fire. Calculate the temperature change vector between two adjacent time steps, as shown in the following formula:

[0030] ;

[0031] To align and concatenate this difference information with the original temperature information in terms of dimension, a difference feature matrix X is constructed. diff (t), as shown in the following formula, is filled with a zero vector because there is no earlier data to calculate the difference at the first time point of the window:

[0032] ;

[0033] Where 0 is a 1×N s A vector of all zeros;

[0034] S124, concatenate the original temperature feature matrix and the temperature difference feature matrix along the column dimension to form a combined feature matrix X. combined(t), as shown in the following formula:

[0035] ;

[0036] Matrix X combined (t) also includes the temperature distribution and its changing trend in the tunnel space at a certain moment, providing rich spatiotemporal information for the multimodal inversion network model.

[0037] According to an embodiment of the present invention, step S13 includes:

[0038] S131, standardize the image size and use bilinear interpolation algorithm to adjust all temperature cloud images to 224×224 pixels;

[0039] S132, standardize the image data, including: first, normalize the pixel values ​​of each channel of the image from the integer range of [0,255] to the floating-point range of [0,1]; second, standardize using the universal mean and standard deviation calculated on the ImageNet dataset; where the mean μ is [0.485,0.456,0.406] and the variance σ is [0.229,0.224,0.225]; after this processing, each channel of the image becomes a distribution with a mean of 0 and a standard deviation of 1.

[0040] According to an embodiment of the present invention, step S14 includes:

[0041] S141, Determine the range of physical parameters, including: fire source location, fire source power, and exhaust fan speed;

[0042] S142. To prevent the model from overfitting due to directly using real fire source information during training, a tiered masking mechanism is adopted for the physical parameter input, specifically including three modes: "none" mode with no physical input, "wind_only" mode with only wind speed, and "full" mode with complete input. Among them, "none" mode is used for strict leak-proof training, "wind_only" mode is used to simulate measurable engineering conditions, and "full" mode is only used in comparative experiments.

[0043] According to an embodiment of the present invention, step S2 includes:

[0044] S21, Construct a time feature extraction module. This module is used to extract the dynamic evolution pattern of fire development from the time-series data of the temperature sensor. The input data for the time feature extraction module is a two-dimensional feature matrix containing the temperature values ​​of the current moment and the previous moment and their first-order difference, with a size of W×2N. sThe temporal feature extraction module employs a bidirectional long short-term memory network structure, consisting of two stacked LSTM layers with 128 hidden units per layer. The bidirectional configuration simultaneously captures forward and backward temporal dependencies, resulting in an output feature dimension of 256. Dropout layers are used between layers to reduce overfitting risk. The activation function combines tanh and sigmoid gating. The output sequence undergoes temporal average pooling to form a temporal feature vector of length 256.

[0045] S22, Construct an image encoding module, which is used to extract spatial distribution features from temperature cloud maps; the image encoding module is implemented based on a pre-trained ResNet-18 convolutional neural network, and the weight parameters are obtained by pre-training on the ImageNet dataset;

[0046] S23, Construct a physical parameter embedding module. The physical parameter embedding module is used to map the fire source location, fire source power and ventilation speed into high-dimensional learnable feature vectors. The physical parameter embedding module consists of two fully connected network layers: the number of channels in the first layer is expanded from 3 to 16, and the number of channels in the second layer is expanded from 16 to 32. The activation function for both layers is ReLU.

[0047] S24, Construct a leak-proof gating fusion module. This module fuses temporal, image, physical, and spatial features and adaptively adjusts the modal weights based on modal availability. The inputs to the leak-proof gating fusion module include temporal features, image features, spatial statistical features, and physical embedding features. The module contains two independent gating networks, which are used to calculate image modal weights and physical modal weights, respectively. Each gating network consists of a linear transform layer and a sigmoid activation function.

[0048] S25, Construct the fire situation inversion head module, which includes a fire source power regression head and a coarse-to-fine fire source location prediction head; among which...

[0049] The input to the fire source power regression head is a 512-dimensional fused feature, which passes through three fully connected network layers in sequence: the first layer's dimension is reduced from 512 to 256, the second layer's dimension is reduced from 256 to 128, and the third layer's dimension is reduced from 128 to 1. The first and second layers use the SiLU activation function and set Dropout to output a single continuous value representing the fire source power. During the training phase, the mean squared error loss function is used for regression optimization.

[0050] The coarse-to-fine fire source location prediction head adopts a two-level structure, including a coarse classification layer, a fine regression layer, and an optional heatmap branch. The coarse classification layer is implemented through two layers of linear mapping: the first layer's dimension is reduced from 512 to 256, and the second layer's dimension is reduced from 256 to 60, with the activation function being ReLU, outputting the logits vector of the fire source interval. The fine regression layer is used to predict continuous offsets within the interval, consisting of two layers of linear mapping, with the output constrained to the range [-0.5, 0.5] by the tanh function and multiplied by the bin width to achieve continuous location prediction. Finally, the predicted fire source location value is jointly determined by the pattern of "coarse classification center + fine regression offset".

[0051] According to an embodiment of the present invention, step S3 includes:

[0052] S31, Constructing a temperature gradient constraint loss, which limits the spatial variation of the inversion results from the multimodal inversion network model, ensuring that the predicted temperature distribution and its derivatives conform to the continuity requirements of the heat conduction equation; specifically including:

[0053] Assume the predicted temperature field is arranged along the longitudinal direction of the tunnel with N s At each measurement point, the multimodal inversion network model outputs a predicted temperature sequence T' at time t. t =[T' t,1 T' t,2 ,…,T' t,Ns The temperature gradient loss is defined as the average of the sum of squares of the temperature differences between adjacent measuring points, as shown in the following formula:

[0054] ;

[0055] S32, Construct the maximum temperature rise constraint loss function. The maximum temperature rise constraint loss is used to constrain the physical correspondence between the fire source power and the peak temperature rise, based on the empirical formula for fire heat release power and temperature rise:

[0056] ;

[0057] Where Q is the fire source power in MW; k is the scaling factor; and Q' is the predicted power output of the multimodal inversion network model at each time step. t With the maximum temperature rise ΔT' max,t This relation is satisfied;

[0058] The loss function in the multimodal inversion network model is shown in the following formula:

[0059] ;

[0060] The Smooth L1 form is adopted, and a lower limit of 0.1 is set in the power input to prevent gradient explosion;

[0061] S33, Construct a longitudinal attenuation constraint loss function. This function constrains the attenuation of predicted temperature along the longitudinal direction of the tunnel, ensuring that temperature predictions in areas far from the fire source conform to the physical characteristics of fire flow. Specifically, this includes:

[0062] Based on the empirical law of temperature decay along the longitudinal direction in tunnel fires, the temperature rise decreases as the distance from the fire source increases. The empirical relationship between temperature rise and distance from the fire source is as follows:

[0063] ;

[0064] Where x is the longitudinal distance from the fire source, W is the tunnel width, T0 is the reference temperature rise value at the current moment, and A, B, and C are empirical parameters;

[0065] During the inversion process, let the fire source location predicted by the multimodal inversion network model at time t be x'. f (t), for the relative distance d of the i-th sensor i (t) is defined as:

[0066] ;

[0067] The temperature or temperature rise at this measuring point predicted by the multimodal inversion network model is denoted as T'. t,i Then, based on the empirical relationship between temperature rise and distance from the fire source, the theoretical attenuation curve is constructed, and the theoretical value T' at that measuring point is calculated. fit,i As shown in the following formula:

[0068] ;

[0069] To constrain the prediction results to follow the longitudinal decay law, a weighted longitudinal decay constraint loss function is defined, as shown in the following formula:

[0070] ;

[0071] In addition, to balance the constraint strength between the near and far regions, a distance weighting function w is introduced. i (t)=e -0.02di(t) This is used to impose higher constraint strength in the area near the fire source and moderately weaken the constraint in the area far from the fire source, so that the overall output of the multimodal inversion network model follows the physical attenuation law;

[0072] S34; During the training phase, an end-to-end joint optimization approach is adopted. In each training iteration, the input consists of multimodal samples containing temporal features, image features, and physical embedding features. The network outputs predicted fire source power, predicted fire source location, and probability distribution of intermediate heat maps. The total loss function is composed of a weighted sum of the main task loss and the physical constraint loss, as shown in the following formula:

[0073] ;

[0074] in, For the total loss function, , , For dynamically adjustable learnable weights, For temperature gradient constrained loss, To minimize temperature rise confinement loss, For longitudinal attenuation constraint losses, the main task losses include the mean square error (MSE) loss of fire source power regression. HRR (t), L1 loss due to fire source location regression L loc (t), coarse classification cross-entropy loss of fire source location L CE (t), the expression is as follows:

[0075] ;

[0076] ;

[0077] ;

[0078] in, To predict power, To represent the actual power of the fire source, To predict the location of the fire source, This is the actual location of the fire source. This represents the predicted probability that the fire source location predicted by the model belongs to the true category;

[0079] The optimization algorithm uses the AdamW optimizer with an initial learning rate of 1×10⁻⁶. -4 Cosine annealing scheduling and early stopping strategies are employed to prevent overfitting and improve convergence stability.

[0080] S35, the training dataset consists of multiple fire condition data. Each dataset includes a time series of temperature sensors, temperature distribution images, and corresponding fire source physical parameters. Before training, all features are standardized: temperature data is normalized using global mean and standard deviation; power data uses an independent normalizer to ensure consistent feature distribution under different conditions. To enhance the model's learning of the early fire stage, a time-weighted sampling strategy is used for the training samples. The sampling weight of early samples within a second preset time period is set to 1.5, and the weight of other samples is 1.0. The second preset time period is shorter than the first preset time period. When using the Weighted Random Sampler sampling method, the sampling probability is proportional to the time weighting value to improve the sensitivity of the multimodal inversion network model to temperature change characteristics in the early stage of fire development.

[0081] S36 introduces a whole-modal Drop strategy during the training phase. Specifically, in each batch of training, the entire image input is set to zero with a probability of 0.35, which is regarded as a camera failure or smoke obstruction. When the image is set to zero, the multimodal inversion network model automatically triggers the "image availability gating signal" to 0, so that the image branch weights are returned to zero, thereby preventing the wrong input of invalid information.

[0082] S37, a high-probability modal Drop operation is performed at the physical input end. The Drop operation sets the entire physical input vector to zero with a probability of 0.7, retaining only non-leaking statistical noise or wind speed information. Specifically: in the "none" mode configuration, the physical input is completely masked, and only the zero vector form is retained; when set to "wind_only" mode, only the wind speed component is retained for ventilation impact modeling; in the "full" mode, it is only used for comparative experiments and is not configured as the main model.

[0083] S38, the sensor random occlusion enhancement mechanism, specifically: in each batch, at most 10% of the sensor channels are randomly selected with a probability of 0.15 and set to zero at all time steps; at the same time, the corresponding temperature difference channels are synchronously set to zero to maintain the consistency of input feature dimensions.

[0084] S39 employs a consistency distillation mechanism to maintain the consistency of model output results under different modal input conditions, thereby improving the stability and generalization performance of the multimodal inversion network model in the case of missing multi-source data. Specifically, in each training process, the multimodal inversion network model is input into the same batch of training samples in two forms: one is a complete modal input, including temperature sensor sequences, temperature images, and physical parameters; the other is a masked modal input, where the images and physical parameters are simultaneously set to zero, retaining only the sensor modality. The former serves as the student model, and the latter as the teacher model, with both calculating the output results in parallel under the same structure.

[0085] S310: During training, various losses are dynamically balanced using uncertainty weighting. These losses include HRR regression, location classification, location regression, physical consistency, and consistency distillation. The optimizer uses the AdamW algorithm with an initial learning rate of 1×10⁻⁶. -4 Weight decay coefficient 1×10 -4 It also uses the ReduceLROnPlateau scheduler based on the validation set loss to automatically adjust the learning rate; mixed precision training is used during training to reduce memory usage, and a gradient clipping cap of 1.0 is set to prevent gradient explosion; when the total validation set loss does not improve for a preset number of consecutive rounds, an early stopping mechanism is triggered to avoid overfitting.

[0086] According to an embodiment of the present invention, step S4 includes:

[0087] S41, Input Data Analysis and Feature Construction: The multimodal inversion network model first reads the temperature sensor data CSV file obtained from fire simulation or on-site collection. The temperature sensor data CSV file contains time series and multiple temperature measurement point records. The multimodal inversion network model automatically identifies the time series and temperature series, and selects equally spaced sensor series for analysis based on the number of sensors during training. After reading, the multimodal inversion network model performs numerical cleaning on all temperature data and generates a time series vector. Then, it constructs a feature tensor according to the sliding window length W=2 during the training phase: at each moment, the feature matrix [W,2S] is composed of the current and previous temperature values ​​and their time differences. The feature mean and standard deviation saved during training are used for normalization to ensure that the input distribution is consistent with the training phase.

[0088] S42, Model Loading and Inference Configuration: The system loads the trained multimodal fusion inversion network model from the specified model directory, including weight files ending in .pth and normalized parameter files ending in .npz; the inference process runs in a frozen state, the parameters of the multimodal fusion inversion network model remain unchanged, and only forward computation is performed; the system supports automatic device selection, running in CUDA mode when an available GPU is detected, otherwise executing on the CPU; image modality and physical modality inputs are disabled during inference, and all-zero tensor inputs are automatically used with gating mechanisms enabled to maintain consistency with leakage prevention during training; sensor coordinates are automatically generated according to tunnel length and spacing for fire source location grid definition and probability calculation;

[0089] S43, Real-time step-by-step inference and result smoothing: The system inputs feature tensors step-by-step and performs forward calculations to obtain the predicted values ​​of fire source power and fire source location. The predicted value of fire source power is output by the fire source power regression head, and the predicted value of fire source location is jointly determined by the coarse classification and fine regression two-level networks.

[0090] During inference, the system retains the hidden states of the LSTM between adjacent time points to maintain temporal continuity. At the same time, in order to suppress high-frequency fluctuations, the system uses an exponential moving average algorithm to smooth the power and position prediction results at consecutive time points, thereby improving curve stability and physical consistency.

[0091] S44, Result File Generation and Visualization Output: After completing the inference, the system automatically outputs the following result files:

[0092] The main results file, in CSV format, contains the timestamp for each moment, the predicted fire source power, the predicted fire source location, and the original prediction results.

[0093] Probability matrix file, NPZ format: stores the probability matrix of fire source location at each time point, the discrete grid coordinates of the tunnel, and the time series;

[0094] Wide table format CSV file: each column corresponds to the tunnel location, and each row is a time point, used to directly generate heat maps;

[0095] Long table format CSV file: contains three columns, namely time, location and probability, for statistical and graphical analysis;

[0096] The fire source power curve plot uses time as the horizontal axis and power as the vertical axis to show the fire development trend; the fire source location probability heat map plots tunnel location as the horizontal axis and probability as the vertical axis to represent the probability of fire source occurrence at each location, intuitively reflecting the spatiotemporal migration process of fire source in the longitudinal direction of the tunnel.

[0097] The tunnel fire inversion method based on multimodal fusion and physical constraints in this invention has the following beneficial results:

[0098] This invention introduces multimodal inputs such as sensor temperature time series, temperature distribution images, and ventilation parameters, combined with a gated fusion network and a leak-proof training mechanism, to achieve joint inversion of fire source power and location. At the same time, physical constraints such as temperature gradient and longitudinal attenuation are introduced into the loss function to ensure that the prediction results conform to thermal laws, thereby improving inversion accuracy and physical consistency while maintaining real-time performance.

[0099] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0100] Figure 1 This is a flowchart of a tunnel fire inversion method based on multimodal fusion and physical constraints according to an embodiment of the present invention;

[0101] Figure 2 This is a framework diagram of a tunnel fire inversion model based on multimodal fusion and physical constraints according to an embodiment of the present invention.

[0102] Figure 3 This is a framework diagram of the various sub-modules of a multimodal inversion network model according to an embodiment of the present invention;

[0103] Figure 4 This is a schematic diagram illustrating the power inversion comparison of different model fire sources under the working conditions of 250m-30MW-3m / s according to an embodiment of the present invention.

[0104] Figure 5 This is a schematic diagram of the probability inversion result of the optimized model fire source location under the working condition of 250m-30MW-3m / s according to an embodiment of the present invention. Detailed Implementation

[0105] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0106] The tunnel fire inversion method based on multimodal fusion and physical constraints proposed in this invention is described below with reference to the accompanying drawings.

[0107] like Figures 1-5 As shown, the tunnel fire inversion method based on multimodal fusion and physical constraints according to an embodiment of the present invention may include the following steps:

[0108] S1 collects time-series data, temperature distribution images, and physical parameters from distributed temperature sensors inside the tunnel to construct a multimodal training dataset.

[0109] According to an embodiment of the present invention, step S1 includes:

[0110] S11 uses FDS simulation to obtain temperature sensor data, temperature cloud images, and physical parameters under different fire source conditions inside the tunnel.

[0111] According to an embodiment of the present invention, step S11 includes:

[0112] S111, construct a tunnel numerical model of a preset size, and determine the number and corresponding location of fire sources, the configuration range of fire source power, and the configuration parameters of longitudinal ventilation in the tunnel numerical model.

[0113] For example, a numerical model of a tunnel with dimensions of 300m (length) × 6m (width) × 9m (height) can be constructed. In this model, the fire source locations are set at three points: 50m, 150m, and 250m from the tunnel entrance; the fire source power is set in 8 levels at 5MW intervals within the range of 5MW to 40MW; and the longitudinal ventilation velocity is set in 12 levels at 1m / s intervals within the range of 1m / s to 12m / s.

[0114] S112, determine the range of the fire source, the power growth of the fire source follows the t² fire model, the total simulation time is the first preset time; the fuel type is heptane, and the carbon monoxide generation rate and particulate matter production rate in the combustion products of the fuel are set respectively.

[0115] For example, the fire source can be set 0.5m above the ground, with a rectangular area of ​​2m × 5m cross-section. Its power increase follows a t² fire model, and the total simulation time is 600 seconds. Heptane (C7H) is used as the fuel type. 16The carbon monoxide generation rate in the combustion products was set to 0.006, and the particulate matter production rate was set to 0.015. The tunnel wall material was defined as concrete, and its thermal parameters were directly adopted from the default configuration of the FDS software.

[0116] S113 employs a dynamic grid partitioning strategy to divide the tunnel space.

[0117] Specifically, to improve computational accuracy and efficiency, a dynamic mesh generation strategy was adopted: a high-precision 0.5m mesh was used for the area within a 50m radius and 5m radius of the fire source center and each jet fan, while a standard 1.0m mesh was used for the remaining areas of the tunnel. The boundary conditions at both ends of the tunnel were set to "OPEN" to simulate an open atmospheric environment.

[0118] S114, the ventilation system includes two sets of jet fans, symmetrically arranged in the tunnel, each set of jet fans contains 2 fans; the volumetric flow rate of the air supply duct is precisely controlled by the HVAC system built into the FDS software to achieve the target wind speed condition.

[0119] For example, the ventilation system consists of two sets of jet fans, symmetrically arranged at distances of 75m and 225m from the tunnel entrance, with each set containing two fans. Each fan measures 3m × 1.5m × 1.5m and is mounted 1m from the tunnel ceiling. The volumetric flow rate of the air supply duct (1m in diameter) is precisely controlled via the FDS's built-in HVAC system to achieve the target air velocity.

[0120] S115: Multiple temperature sensors are deployed at preset intervals along the longitudinal centerline of the tunnel. Each sensor records temperature data, with a sampling interval of 1 second. Simultaneously, a temperature monitoring plane is set at the longitudinal center of the tunnel to export temperature cloud map data, with an image sampling interval of 1 second. For example, in a 300-meter-long tunnel model, temperature sensors are deployed at 5-meter intervals along the longitudinal centerline of the tunnel, for a total of 60 sensors.

[0121] S12 performs sliding window processing on the temperature sensor data to extract the current temperature and its first-order difference features, forming a time-series feature matrix.

[0122] According to an embodiment of the present invention, step S12 includes:

[0123] S121, Set a sliding window to extract features from the time series data of each sensor using a sliding window of fixed length.

[0124] For example, setting the window length to W=2 and the step size to S=1 means that the window moves forward one time step at a time on the time axis. This method ensures that adjacent samples overlap, thus preserving the continuity of temperature changes and the temporal correlation of the data.

[0125] S122, for any time point t, extract the matrix of temperature readings from all sensors at the current time and the previous time, as shown in the following formula:

[0126] ;

[0127] Among them, T i (t) represents the temperature value at the i-th measuring point at time t, N s The number of sensors; the temperature feature matrix X at time t. temp (t) is a W-row N-row s A column matrix, expressed as follows:

[0128] ;

[0129] The first row of the matrix represents the global temperature at the previous time t−1, and the second row represents the global temperature at the current time t.

[0130] S123, calculate the temperature difference characteristics. The instantaneous rate of temperature change (approximately the first-order difference) contains important information about the dynamic evolution of the fire. Calculate the temperature change vector between two adjacent time steps, as shown in the following formula:

[0131] ;

[0132] To align and concatenate this difference information with the original temperature information in terms of dimension, a difference feature matrix X is constructed. diff (t), as shown in the following formula, is filled with a zero vector because there is no earlier data to calculate the difference at the first time point of the window:

[0133] ;

[0134] Where 0 is a 1×N s A vector of all zeros.

[0135] S124, concatenate the original temperature feature matrix and the temperature difference feature matrix along the column dimension to form a combined feature matrix X. combined (t), as shown in the following formula:

[0136] ;

[0137] Matrix X combined (t) also includes the temperature distribution and its changing trend in the tunnel space at a certain moment, providing rich spatiotemporal information for the multimodal inversion network model.

[0138] S13 performs standardization and size unification processing on the temperature distribution image.

[0139] According to an embodiment of the present invention, step S13 includes:

[0140] S131, standardize the image size and use bilinear interpolation algorithm to adjust all temperature cloud images to 224×224 pixels;

[0141] S132, standardize the image data, including: first, normalize the pixel values ​​of each channel of the image from the integer range of [0,255] to the floating-point range of [0,1]; second, standardize using the universal mean and standard deviation calculated on the ImageNet dataset; where the mean μ is [0.485,0.456,0.406] and the variance σ is [0.229,0.224,0.225]; after this processing, each channel of the image becomes a distribution with a mean of 0 and a standard deviation of 1.

[0142] S14, Physical parameter normalization and leakage prevention measures.

[0143] According to an embodiment of the present invention, step S14 includes:

[0144] S141, Determine the range of physical parameters, including: fire source location, fire source power, and exhaust fan speed.

[0145] For example, the location of the fire source L fire Within a 300-meter-long tunnel, its value ranges from 0 to 300m; ignition source power P fire Based on the design conditions, its value ranges from 5 to 40 MW; the exhaust fan velocity V wind According to the working conditions, its value range is 1~12m / s.

[0146] Furthermore, to eliminate the impact of differences in the dimensions and numerical ranges of different physical parameters on model training, Z-score standardization (also known as standard deviation standardization) is performed on each parameter to make it conform to a standard normal distribution with a mean of 0 and a standard deviation of 1. For each physical parameter in the training dataset, its mean μ and standard deviation σ are calculated, and then transformed as shown in the following formula:

[0147] ;

[0148] ;

[0149] ;

[0150] In the formula, μ L μ P μ VThese are the arithmetic mean of the fire source location, fire source power, and ventilation speed for all samples in the training dataset; σ L , σ P , σ V These are the standard deviations of the fire source location, fire source power, and ventilation speed for all samples in the training dataset.

[0151] S142. To prevent the model from overfitting due to directly using real fire source information during training, a tiered masking mechanism is adopted for the physical parameter input, specifically including three modes: "none" mode with no physical input, "wind_only" mode with only wind speed, and "full" mode with complete input. Among them, "none" mode is used for strict leak-proof training, "wind_only" mode is used to simulate measurable engineering conditions, and "full" mode is only used in comparative experiments.

[0152] Specifically, the "none" mode, which has no physical input, masks all physical parameter inputs, and its feature vector is p=[0,0,0].

[0153] Wind-only mode: This mode retains only observable ventilation wind speed input, with feature vector p=[0,0,V]. norm ];

[0154] Full input mode: This mode simultaneously inputs the fire source location, power, and wind speed, constructing the feature vector as p=[L norm , P norm V norm ].

[0155] The "none" mode is used for rigorous leak-proof training, the "wind_only" mode is used to simulate engineering-measurable conditions, and the "full" mode is used only in comparative experiments. This hierarchical mechanism ensures that the model maintains stable generalization performance under different information availability scenarios.

[0156] S2 establishes a multimodal inversion network model that includes temporal feature extraction, image encoding, physical parameter embedding, and gating fusion. The sub-modules of the multimodal inversion network model are as follows: Figure 3 As shown.

[0157] According to an embodiment of the present invention, step S2 includes:

[0158] S21, Construct a time feature extraction module. This module is used to extract the dynamic evolution pattern of fire development from the time-series data of the temperature sensor. The input data for the time feature extraction module is a two-dimensional feature matrix containing the temperature values ​​of the current moment and the previous moment and their first-order difference, with a size of W×2N. sThe temporal feature extraction module employs a bidirectional long short-term memory (LSTM) network structure, comprising two stacked LSTM layers with 128 hidden units per layer. This bidirectional configuration simultaneously captures forward and backward temporal dependencies, resulting in an output feature dimension of 256. Dropout layers are used between layers to reduce overfitting risk. The activation function combines tanh and sigmoid gating. The output sequence undergoes temporal average pooling to form a 256-dimensional temporal feature vector. This module effectively extracts the phased features of fire development and temperature change trends, providing a temporal information foundation for subsequent multimodal feature fusion.

[0159] S22, Construct an image encoding module, which is used to extract spatial distribution features from temperature cloud maps; the image encoding module is implemented based on a pre-trained ResNet-18 convolutional neural network, and the weight parameters are obtained by pre-training on the ImageNet dataset.

[0160] Specifically, to balance accuracy and computational efficiency, the network structure is retained up to the penultimate layer (removing the fully connected classification layer and the average pooling layer), with an output feature dimension of [512, H', W'], where H' and W' are the spatial dimensions of the feature map. Subsequently, the feature map is compressed to [512, 1, 1] through an adaptive average pooling layer (AdaptiveAvgPool2d), and then converted into a 256-dimensional feature vector through a flattening layer and a linear mapping layer (Linear). The activation function is ReLU, followed by a dropout layer with a dropout rate of 0.3. To improve the robustness of the model, random modality dropout (probability 0.35) is introduced at the input during the training phase, i.e., the entire image is zeroed with a certain probability. When an image modality is missing, the system automatically inputs a zero tensor and automatically adjusts the feature weights through a gating mechanism to ensure that the network structure is consistent between the training and inference phases.

[0161] S23, Construct a physical parameter embedding module. The physical parameter embedding module is used to map the fire source location, fire source power and ventilation speed into high-dimensional learnable feature vectors. The physical parameter embedding module consists of two fully connected network layers: the number of channels in the first layer is expanded from 3 to 16, and the number of channels in the second layer is expanded from 16 to 32. The activation function for both layers is ReLU.

[0162] Specifically, the input vector is configured according to the "physics_input_mode" parameter and can use three modes:

[0163] (1) None mode: Disables all physical input;

[0164] (2) wind_only mode: Input only the wind speed;

[0165] (3) Full mode: Full input of location, power and wind speed (for comparative experiments only).

[0166] To further enhance the model's generalization ability, a modality discarding operation is applied to the physical input with a probability of 0.7 during the training phase, which randomly sets the entire physical vector to zero. The final output is a 32-dimensional embedded feature, used for subsequent multimodal fusion.

[0167] S24. Construct a leak-proof gating fusion module. This module fuses temporal, image, physical, and spatial features and adaptively adjusts the weights of each modality based on modality availability. The inputs to the leak-proof gating fusion module include: temporal features (256 dimensions), image features (256 dimensions), spatial statistical features (extracted by a one-dimensional convolutional module, 64 dimensions), and physical embedding features (32 dimensions). The leak-proof gating fusion module sets up two independent gating networks to calculate image modality weights and physical modality weights, respectively. Each gating network consists of a linear transformation layer and a sigmoid activation function.

[0168] The specific settings for the two gating systems are as follows:

[0169] (1) The image gating input is 256+32+64+1 dimensional, where “1” is the image validity flag;

[0170] (2) The physical gating input is 256+32+64+1 dimensional, where “1” is the physical modality availability flag;

[0171] Two types of gated outputs are multiplied element-wise with their corresponding modal features to achieve adaptive modality adjustment. The weighted modal features are concatenated along the channel dimension and then input into a linear fusion layer (512 dimensions) to generate a fused feature vector. A dropout layer (with a dropout rate of 0.4) is used to prevent feature-dominated shifts. This module effectively avoids the risk of information leakage while maintaining multimodal complementarity, enabling the network to maintain stable output even in cases of missing or abnormal modalities.

[0172] S25, Construct the fire situation inversion head module, which includes a fire source power regression head and a coarse-to-fine fire source location prediction head; among which...

[0173] The input to the fire source power regression head is a 512-dimensional fused feature, which passes through three fully connected network layers in sequence: the first layer reduces the dimension from 512 to 256, the second layer reduces the dimension from 256 to 128, and the third layer reduces the dimension from 128 to 1. The first and second layers use the SiLU activation function and set Dropout (dropout rate 0.3), and output a single continuous value representing the fire source power. The mean squared error (MSE) loss function is used for regression optimization during the training phase.

[0174] The coarse-to-fine fire source location prediction head adopts a two-level structure, including a coarse classification layer, a fine regression layer, and an optional heatmap branch. The 300m longitudinal range of the tunnel is divided into equidistant intervals, each 5m wide, with a total of 60 categories. The coarse classification layer is implemented through two linear mapping layers: the first layer's dimension is reduced from 512 to 256, and the second layer's dimension is reduced from 256 to 60, with ReLU activation, outputting the logits vector of the fire source interval. The fine regression layer is used to predict continuous offsets within the interval, consisting of two linear mapping layers. The output is constrained to the range [-0.5, 0.5] by the tanh function and multiplied by the bin width to achieve continuous location prediction. When the configuration option `use_loc_heatmap_aux=True`, the additional heatmap branch generates a continuous location grid of length G=301, and calculates the weighted positions using SoftArgmax to improve spatial smoothness.

[0175] The final predicted fire source location is determined by a combination of coarse classification centers and fine regression bias. During the training phase, classification cross-entropy loss, location L1 loss, and physical consistency constraint loss are used in combination, and multiple task objectives are dynamically balanced through uncertainty weighting.

[0176] S3. Design a physical constraint loss and information leakage prevention mechanism, and train the multimodal inversion network model using a multimodal training dataset.

[0177] According to an embodiment of the present invention, step S3 includes:

[0178] S31, Constructing a temperature gradient constraint loss, which limits the spatial variation of the inversion results from the multimodal inversion network model, ensuring that the predicted temperature distribution and its derivatives conform to the continuity requirements of the heat conduction equation; specifically including:

[0179] Assume the predicted temperature field is arranged along the longitudinal direction of the tunnel with N s At each measurement point, the multimodal inversion network model outputs a predicted temperature sequence T' at time t. t =[T' t,1 T' t,2 ,…,T' t,Ns The temperature gradient loss is defined as the average of the sum of squares of the temperature differences between adjacent measuring points, as shown in the following formula:

[0180] ;

[0181] This constraint term ensures that the model outputs a smooth and continuous temperature distribution in the longitudinal spatial direction, preventing unreasonable and drastic fluctuations. To avoid excessively strong gradients near the fire source that could cause constraint imbalance, a weight attenuation coefficient of 0.5 is applied to three measurement points in the vicinity of the fire source.

[0182] S32, Construct the maximum temperature rise constraint loss function. The maximum temperature rise constraint loss is used to constrain the physical correspondence between the fire source power and the peak temperature rise, based on the empirical formula for fire heat release power and temperature rise:

[0183] ;

[0184] Where Q is the fire source power in MW; k is the scaling factor; and Q' is the predicted power output of the multimodal inversion network model at each time step. t With the maximum temperature rise ΔT' max,t This relation is satisfied;

[0185] The loss function in the multimodal inversion network model is shown in the following formula:

[0186] ;

[0187] The Smooth L1 form is adopted, and a lower limit of 0.1 is set in the power input to prevent gradient explosion;

[0188] This constraint ensures that the temperature rise prediction is not saturated under high power conditions and that the temperature rise is not overestimated under low power conditions, thereby maintaining the consistency of the heat source thermal characteristics.

[0189] S33, Construct a longitudinal attenuation constraint loss function. This function constrains the attenuation of predicted temperature along the longitudinal direction of the tunnel, ensuring that temperature predictions in areas far from the fire source conform to the physical characteristics of fire flow. Specifically, this includes:

[0190] Based on the empirical law of temperature decay along the longitudinal direction in tunnel fires, the temperature rise decreases as the distance from the fire source increases. The empirical relationship between temperature rise and distance from the fire source is as follows:

[0191] ;

[0192] Where x is the longitudinal distance from the fire source, W is the tunnel width, T0 is the reference temperature rise value at the current moment, and A, B, and C are empirical parameters;

[0193] During the inversion process, let the fire source location predicted by the multimodal inversion network model at time t be x'. f (t), for the relative distance d of the i-th sensor i (t) is defined as:

[0194] ;

[0195] The temperature or temperature rise at this measuring point predicted by the multimodal inversion network model is denoted as T'. t,i Then, based on the empirical relationship between temperature rise and distance from the fire source, the theoretical attenuation curve is constructed, and the theoretical value T' at that measuring point is calculated. fit,iAs shown in the following formula:

[0196] ;

[0197] To constrain the prediction results to follow the longitudinal decay law, a weighted longitudinal decay constraint loss function is defined, as shown in the following formula:

[0198] ;

[0199] In addition, to balance the constraint strength between the near and far regions, a distance weighting function w is introduced. i (t)=e -0.02di(t) This is used to impose higher constraint strength in the area near the fire source and moderately weaken the constraint in the area far from the fire source, so that the overall output of the multimodal inversion network model follows the physical attenuation law.

[0200] S34; During the training phase, an end-to-end joint optimization approach is adopted. In each training iteration, the input consists of multimodal samples containing temporal features, image features, and physical embedding features. The network outputs predicted fire source power, predicted fire source location, and probability distribution of intermediate heat maps. The total loss function is composed of a weighted sum of the main task loss and the physical constraint loss, as shown in the following formula:

[0201] ;

[0202] in, For the total loss function, , , For dynamically adjustable learnable weights, For temperature gradient constrained loss, To minimize temperature rise confinement loss, For longitudinal attenuation constraint losses, the main task losses include the mean square error (MSE) loss of fire source power regression. HRR (t), L1 loss due to fire source location regression L loc (t), coarse classification cross-entropy loss of fire source location L CE (t), the expression is as follows:

[0203] ;

[0204] ;

[0205] ;

[0206] in, To predict power, To represent the actual power of the fire source, To predict the location of the fire source, This is the actual location of the fire source. This represents the predicted probability that the fire source location predicted by the model belongs to the true category;

[0207] The optimization algorithm uses the AdamW optimizer with an initial learning rate of 1×10⁻⁶. -4 Cosine annealing scheduling and early stopping strategies are employed to prevent overfitting and improve convergence stability.

[0208] S35. To avoid the model's over-reliance on visual modalities during training, which could lead to physical leakage risks during inference, a whole-modality drop strategy is introduced during training. Specifically: the training dataset consists of multiple fire condition data sets, each including temperature sensor time-series sequences, temperature distribution images, and corresponding fire source physical parameters; before training, all features are standardized: temperature data is normalized using global mean and standard deviation; power data uses independent normalizers to ensure consistent feature distribution across different conditions; to enhance the model's learning of early fire stages, a time-weighted sampling strategy is used for training samples, with the sampling weight of early samples within a second preset time period set to 1.5, and the remaining samples to 1.0. The second preset time period is shorter than the first preset time period; for example, when the first preset time period is 600s, the second preset time period is 120s; when using the Weighted Random Sampler method, the sampling probability is proportional to the time-weighted weights to improve the sensitivity of the multimodal inversion network model to temperature change characteristics in the early stages of fire development.

[0209] S36. To avoid the model's over-reliance on the visual modality during training, which could lead to physical leakage during inference, a whole-modality drop strategy is introduced during training. Specifically, in each batch of training, the entire image input is zeroed with a probability of 0.35, considered as a camera malfunction or smoke occlusion condition. When the image is zeroed, the multimodal inversion network model automatically triggers the "image availability gating signal" to 0, causing the image branch weights to return to zero, thereby preventing the input of invalid information. Through this mechanism, the model can maintain stable inversion performance even when images are missing or abnormal, ensuring that the training process does not rely on any image channels that may contain ground truth information.

[0210] S37. To further prevent implicit leakage of data during the training phase by utilizing real fire source location, power, and other labeled information, a high-probability modal Drop operation is implemented at the physical input. This Drop operation sets the entire physical input vector to zero with a probability of 0.7, retaining only non-leaking statistical noise or wind speed information. Specifically: in the "none" mode configuration, the physical input is completely masked, retaining only the zero vector form; in the "wind_only" mode, only the wind speed component is retained for ventilation impact modeling; in the "full" mode, it is only used for comparative experiments and not as the main model configuration. This anti-leakage strategy completely eliminates the risk of explicit truth injection without compromising physical consistency, fundamentally ensuring the generalizability and engineering security of the network training process.

[0211] S38. To improve the model's robustness to sensor failures, signal noise, and channel loss, a random occlusion mechanism is implemented during training to enhance sensor random occlusion. Specifically, in each batch, up to 10% of sensor channels are randomly selected with a probability of 0.15 and zeroed out at all time steps; simultaneously, the corresponding temperature difference channels are also zeroed out to maintain consistency in the input feature dimensions. This mechanism significantly improves the model's self-recovery capability in real-world fire monitoring scenarios, particularly in the face of sensor anomalies and short-term disconnections.

[0212] S39 employs a consistency distillation mechanism to maintain the consistency of model output results under different modal input conditions, thereby improving the stability and generalization performance of the multimodal inversion network model in the case of missing multi-source data. Specifically, in each training process, the multimodal inversion network model is input into the same batch of training samples in two forms: one is a complete modal input, including temperature sensor sequences, temperature images, and physical parameters; the other is a masked modal input, in which the images and physical parameters are simultaneously set to zero, retaining only the sensor modality. The former serves as the student model, and the latter serves as the teacher model, with both calculating the output results in parallel under the same structure.

[0213] S310: During training, various losses are dynamically balanced using uncertainty weighting. These losses include HRR regression, location classification, location regression, physical consistency, and consistency distillation. The optimizer uses the AdamW algorithm with an initial learning rate of 1×10⁻⁶. -4 Weight decay coefficient 1×10 -4 It also uses the ReduceLROnPlateau scheduler based on the validation set loss to automatically adjust the learning rate; during training, mixed precision training is used to reduce memory usage, and a gradient clipping cap of 1.0 is set to prevent gradient explosion; when the total validation set loss does not improve for a preset number of rounds (e.g., 8 rounds), an early stopping mechanism is triggered to avoid overfitting.

[0214] S4 utilizes the trained multimodal inversion network model to invert the power and location of the fire source in real time and outputs the fire inversion results.

[0215] According to an embodiment of the present invention, step S4 includes:

[0216] S41, Input Data Analysis and Feature Construction: The multimodal inversion network model first reads the temperature sensor data CSV file obtained from fire simulation or on-site collection. The temperature sensor data CSV file contains time series and multiple temperature measurement point records. The multimodal inversion network model automatically identifies the time series and temperature series, and selects equally spaced sensor series for analysis based on the number of sensors during training. After reading, the multimodal inversion network model performs numerical cleaning on all temperature data and generates a time series vector. Then, it constructs a feature tensor according to the sliding window length W=2 during the training phase: at each moment, the feature matrix [W,2S] is composed of the current and previous temperature values ​​and their time differences. The feature mean and standard deviation saved during training are used for normalization to ensure that the input distribution is consistent with the training phase.

[0217] S42, Model Loading and Inference Configuration: The system loads the trained multimodal fusion inversion network model from the specified model directory, including weight files ending in .pth and normalized parameter files ending in .npz; the inference process runs in a frozen state, the parameters of the multimodal fusion inversion network model remain unchanged, and only forward computation is performed; the system supports automatic device selection, running in CUDA mode when an available GPU is detected, otherwise executing on the CPU; image modality and physical modality inputs are disabled during inference, and all-zero tensor inputs are automatically used with gating mechanisms enabled to maintain consistency with leakage prevention during training; sensor coordinates are automatically generated according to tunnel length and spacing for fire source location grid definition and probability calculation;

[0218] S43, Real-time step-by-step inference and result smoothing: The system inputs feature tensors step-by-step and performs forward calculations to obtain the predicted values ​​of fire source power and fire source location. The predicted value of fire source power is output by the fire source power regression head, and the predicted value of fire source location is jointly determined by the coarse classification and fine regression two-level networks.

[0219] During inference, the system retains the hidden states of the LSTM between adjacent time points to maintain temporal continuity. At the same time, in order to suppress high-frequency fluctuations, the system uses an exponential moving average algorithm to smooth the power and position prediction results at consecutive time points, thereby improving curve stability and physical consistency.

[0220] S44, Result File Generation and Visualization Output: After completing the inference, the system automatically outputs the following result files:

[0221] The main results file, in CSV format, contains the timestamp for each moment, the predicted fire source power, the predicted fire source location, and the original prediction results.

[0222] Probability matrix file, NPZ format: stores the probability matrix of fire source location at each time point, the discrete grid coordinates of the tunnel, and the time series;

[0223] Wide table format CSV file: each column corresponds to the tunnel location, and each row is a time point, used to directly generate heat maps;

[0224] Long table format CSV file: contains three columns, namely time, location and probability, for statistical and graphical analysis;

[0225] The fire source power curve plot uses time as the horizontal axis and power as the vertical axis to show the fire development trend; the fire source location probability heat map plots tunnel location as the horizontal axis and probability as the vertical axis to represent the probability of fire source occurrence at each location, intuitively reflecting the spatiotemporal migration process of fire source in the longitudinal direction of the tunnel.

[0226] like Figure 4 As shown in the figure, this diagram illustrates the difference between the model-predicted fire source power and the actual fire source power at different times; Figure 5 As shown in the figure, this figure illustrates the probability of the fire source location predicted by the model within the first 5 seconds.

[0227] In summary, the tunnel fire inversion method based on multimodal fusion and physical constraints proposed in this invention aims to solve the problem of difficulty in accurately and in real-time inverting the location and power of fire sources in traditional fire monitoring. This method constructs a spatiotemporal feature fusion network by collecting distributed temperature sensor data, temperature distribution images, and physical parameters within the tunnel; it designs an information leakage prevention training mechanism, employing modal random discarding and sensor shielding techniques to avoid overfitting the model to specific input patterns; it introduces a physical constraint loss function, combining temperature gradient, maximum temperature rise, and longitudinal attenuation laws to ensure the physical rationality of the inversion results; and it adopts a coarse-fine two-level location prediction architecture to achieve meter-level accuracy in locating the fire source. This invention can accurately and invert the power and location of fire sources in real time, providing reliable technical support for intelligent monitoring and emergency rescue of tunnel fires.

[0228] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0229] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0230] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components, unless otherwise explicitly limited. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0231] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A tunnel fire inversion method based on multi-modal fusion and physical constraints, characterized in that, The method comprises: S1, collecting distributed temperature sensor time series data, temperature distribution images and physical parameters in the tunnel to construct a multi-modal training data set; S2, establishing a multi-modal inversion network model comprising time feature extraction, image encoding, physical parameter embedding and gate-controlled fusion; wherein step S2 comprises: S21, constructing a time feature extraction module; S22, constructing an image encoding module; S23, constructing a physical parameter embedding module; S24, constructing a leak-proof gate-controlled fusion module, which is used to fuse time, image, physical and spatial features and to adaptively adjust the weights of each modality according to the availability of the modalities; the input of the leak-proof gate-controlled fusion module comprises time features, image features, spatial statistical features and physical embedding features; two independent gate networks are arranged in the leak-proof gate-controlled fusion module, which are respectively used to calculate the image modality weight and the physical modality weight, and each gate network is composed of a linear transformation layer and a sigmoid activation function; S25, constructing a fire condition inversion head module; S3, designing a physical constraint loss and an anti-information leakage mechanism, and training the multi-modal inversion network model using the multi-modal training data set; wherein step S3 comprises: S31, constructing a temperature gradient constraint loss, which is used to limit the variation amplitude of the inversion result of the multi-modal inversion network model in the spatial dimension, so that the predicted temperature distribution and its derivative meet the continuity requirements of the heat conduction equation; S32, constructing a maximum temperature rise constraint loss function, which is used to constrain the physical correspondence between the fire source power and the temperature rise peak value; S33, constructing a longitudinal attenuation constraint loss function, which is used to constrain the attenuation law of the predicted temperature along the longitudinal direction of the tunnel, and to ensure that the temperature prediction in the area far from the fire source meets the physical characteristics of fire flow; S4, using the trained multi-modal inversion network model to real-time invert the fire source power and position, and outputting the fire condition inversion result.

2. The tunnel fire inversion method based on multi-modal fusion and physical constraints according to claim 1, characterized in that, Step S1 comprises: S11, obtaining temperature sensor data, temperature cloud image and physical parameters under different fire source conditions in the tunnel through FDS simulation; S12, performing sliding window processing on the temperature sensor data to extract the current time temperature and its first-order differential features to form a time series feature matrix; S13, performing standardization and size unification processing on the temperature distribution image; S14, performing normalization and leak-proof processing on the physical parameters.

3. The tunnel fire inversion method based on multi-modal fusion and physical constraints according to claim 2, characterized in that, Step S11 comprises: S111, constructing a tunnel numerical model of a preset size, and determining the number and corresponding position of the fire sources, the configuration range of the fire source power, and the configuration parameters of the longitudinal ventilation in the tunnel numerical model; S112, determining the range of the fire sources, and the power growth of the fire sources follows the t² fire model, and the total simulation time is a first preset time; the fuel type is heptane, and the carbon monoxide generation rate and smoke particle yield of the combustion products of the fuel are respectively set; S113, dividing the tunnel space by using a dynamic grid division strategy; S114, the ventilation system includes two groups of jet fans, symmetrically arranged in the tunnel, each group of jet fans contains 2 fans; the volume flow of the air supply pipeline is accurately controlled by the HVAC system built in the FDS software to realize the target wind speed working condition; S115, along the longitudinal center line of the tunnel, a plurality of temperature sensors are arranged at a preset interval, each sensor records temperature data, the sampling time interval is 1 second, and a temperature monitoring plane is arranged at the longitudinal center position of the tunnel, and temperature cloud data is derived, and the image sampling time interval is 1 second.

4. The tunnel fire inversion method based on multi-modal fusion and physical constraints according to claim 2, characterized in that, Step S12 includes: S121, setting a sliding window, and extracting features of time series data of each sensor by using a sliding window with a fixed length; S122, for any time point t, the temperature readings of all sensors at the current time and the previous time are extracted, as shown in the following formula: ; Wherein, T i (t) represents the temperature value at the i-th measuring point position at time t, N s is the number of sensors; the temperature feature matrix X temp (t) corresponding to time t is a W-row N s column matrix, and the expression is as follows: ; Wherein, the first row of the matrix is the global temperature at the previous time t-1, and the second row is the global temperature at the current time t; S123, calculate the temperature difference feature, the instantaneous change rate of temperature contains important information of fire dynamic evolution, calculate the temperature change vector between adjacent two time steps, as shown in the following formula: ; To align and concatenate this difference information with the original temperature information in dimension, a difference feature matrix X diff (t) is filled with a zero vector as there is no earlier data to compute the difference at the first time point of the window: ; where 0 is a 1 x N vector of all zeros s ; and S124, splicing the original temperature feature matrix and the temperature difference feature matrix in the column dimension to form a combined feature matrix X combined (t), as shown in the following formula: ; Matrix X combined (t) contains both the temperature distribution of the tunnel space at a certain moment and its change trend, providing rich spatiotemporal information for the multi-modal inversion network model.

5. The tunnel fire inversion method based on multi-modal fusion and physical constraints according to claim 2, characterized in that, Step S13 includes: S131, standardize the image size, and use a bilinear interpolation algorithm to adjust all temperature cloud images to 224x224 pixels; S132, standardize the image data, including: first, normalize the pixel value of each channel of the image from the integer range [0, 255] to the floating point number range [0, 1]; second, use the general mean and standard deviation calculated on the ImageNet dataset for standardization; wherein, the mean μ is [0.485, 0.456, 0.406], and the variance σ is [0.229, 0.224, 0.225]; after this processing, each channel of the image becomes a distribution with a mean of 0 and a standard deviation of 1.

6. The tunnel fire inversion method based on multi-modal fusion and physical constraints according to claim 2, characterized in that, Step S14 includes: S141, determine the range of physical parameters, wherein the physical parameters include: fire source position, fire source power, and exhaust fan speed; S142, to prevent overfitting caused by directly using real fire source information during model training, a hierarchical shielding mechanism is adopted for physical parameter input, including three modes: "none" mode without physical input mode, "wind only" mode with only wind speed mode, and "full" mode with complete input mode; wherein, the "none" mode is used for strict leak-proof training, the "wind only" mode is used for simulating measurable conditions in engineering, and the "full" mode is only used in comparative experiments.

7. The tunnel fire inversion method based on multi-modal fusion and physical constraints according to claim 1, characterized in that, In S21, the time feature extraction module is configured to extract the dynamic evolution law of fire development from the time series data of the temperature sensor; the input data of the time feature extraction module is a two-dimensional feature matrix containing the temperature values at the current time and the previous time and the first-order difference thereof, and the size is Wx2N s ; the time feature extraction module adopts a bidirectional long short-term memory network structure, and contains two layers of stacked LSTM units, the number of hidden units of each layer is 128, the bidirectional configuration simultaneously captures the forward and backward time dependencies, and the output feature dimension is 256; a Dropout layer is arranged between each layer to reduce the risk of overfitting; the activation function adopts a combination of tanh and sigmoid gate; after the output sequence is subjected to average pooling in the time dimension, a time feature vector with a length of 256 is formed; In S22, the image encoding module is used to extract spatial distribution features from the temperature cloud image; the image encoding module is realized based on a pre-trained ResNet-18 convolutional neural network, and the weight parameters are pre-trained on the ImageNet dataset; In S23, the physical parameter embedding module is configured to map the fire source position, the fire source power and the ventilation wind speed into a high-dimensional learnable feature vector; The physical parameter embedding module is composed of two fully connected networks: the first layer expands the channel number from 3 to 16, and the second layer expands from 16 to 32, and the activation function is ReLU; In S25, the fire situation inversion head module includes a fire source power regression head and a fire source position coarse-precision prediction head; wherein, The input of the fire source power regression head is a 512-dimensional fusion feature, which sequentially passes through three fully connected networks: the first layer reduces the dimension from 512 to 256, the second layer reduces the dimension from 256 to 128, and the third layer reduces the dimension from 128 to 1; the first layer and the second layer adopt SiLU activation function and set Dropout; a single continuous value is output, representing the fire source power; the mean square error loss function is used for regression optimization in the training stage; The fire source position coarse-precision prediction head adopts a Coarse-to-Fine two-level structure, including a coarse classification layer, a fine regression layer and an optional heat map branch; the coarse classification layer is realized by two linear mappings: the first layer reduces the dimension from 512 to 256, and the second layer reduces the dimension from 256 to 60; the activation function is ReLU; the logits vector of the fire source interval is output; the fine regression layer is used to predict the continuous offset in the interval and is composed of two linear mappings; the output is constrained in the range of [-0.5, 0.5] by the tanh function and multiplied by the bin width to realize continuous position prediction; finally, the fire source position prediction value is determined by the mode of "coarse classification center + fine regression offset".

8. The tunnel fire inversion method based on multi-modal fusion and physical constraints according to claim 1, characterized in that, Step S3 further includes: S31 specifically includes: Assume that the predicted temperature field is arranged with N s measurement points along the longitudinal direction of the tunnel, and the predicted temperature sequence output by the multi-modal inversion network model at time t is T' t =[T' t,1 , T' t,2 , …, T' t,Ns ], and the temperature gradient loss is defined as the average of the square sum of the temperature difference between adjacent measurement points, as shown in the following formula: ; S32 specifically includes: according to the empirical formula of fire heat release power and temperature rise: ; Wherein, Q is the fire source power, unit is MW; k is the proportional coefficient; the prediction power Q' of each time output by the multi-modal inversion network model t Satisfies the empirical formula; ΔT' max,t Satisfies the empirical formula; ΔT' The loss function in the multi-modal inversion network model is shown in the following formula: ; Smooth L1 form is adopted, and a lower limit of 0.1 is set in the power input to prevent gradient explosion; S33 specifically includes: According to the empirical decay law of tunnel fire temperature along the longitudinal direction, the temperature rise decays with the increase of the distance from the fire source, and the empirical relationship between the temperature rise and the distance from the fire source is as follows: ; Wherein, x is the longitudinal distance from the fire source, W is the tunnel width, T0 is the reference temperature rise amplitude at the current time, A, B, C are empirical parameters; In the inversion process, record the multimodal inversion network model at time t fire location prediction as x' f (t), the relative distance d of the i-th sensor i (t) is defined as: ; The temperature or temperature rise of the measuring point predicted by the multi-modal inversion network model is denoted as T' t,i According to the empirical relationship between the temperature rise and the distance of the fire source, a theoretical value T' of the theoretical attenuation curve at the measuring point is constructed fit,i As shown in the following formula: ; In order to constrain the prediction result to follow the longitudinal decay law, a weighted longitudinal decay constraint loss function is defined, as shown in the following formula: ; In addition, in order to balance the constraint strength of the near zone and the far zone, a distance weight function w i (t)=e -0.02di(t) is introduced, which is used to give higher constraint strength in the near fire source area and appropriately weaken the constraint in the area far away from the fire source, so as to make the multi-modal inversion network model output overall follow the physical attenuation law. S34; In the training stage, an end-to-end joint optimization method is adopted; in each training iteration, a multi-modal sample containing time features, image features and physical embedding features is input, and the network outputs the fire source power prediction value, the fire source position prediction value and the intermediate heat map probability distribution; the total loss function is composed of the weighted sum of the main task loss and the physical constraint loss, as shown in the following formula: ; wherein, is a total loss function, , , is a dynamically adjusted learnable weight, is a temperature gradient constraint loss, is a maximum temperature rise constraint loss, is a longitudinal decay constraint loss, the main task loss includes a fire power regression mean square error (MSE) loss L HRR (t), a fire location regression L1 loss L loc (t), a fire location coarse classification cross-entropy loss L CE (t), and the expression is as follows: ; ; ; wherein, is a predicted fire source power, is a true fire source power, is a predicted fire source location, is a true fire source location, is a predicted probability that the fire source location predicted by the model belongs to the true class; The optimization algorithm adopts AdamW optimizer, and the initial learning rate is 1x10 -4 Cosine annealing scheduling and early stopping strategy are adopted to prevent overfitting and improve convergence stability. S35, the training data set is composed of multiple fire working condition data, each set of data including a temperature sensor time sequence, a temperature distribution image and corresponding fire source physical parameters; before training, all features are standardized: temperature data is normalized by global mean and standard deviation; power data is standardized independently to ensure consistent feature distribution under different working conditions; to strengthen the model's learning of the early fire stage, the training samples use a time-weighted sampling strategy, with a sampling weight of 1.5 for early samples in a second preset time and 1.0 for the rest, wherein the second preset time is less than the first preset time; when using the Weighted Random Sampler sampling method, the sampling probability is proportional to the time-weighted weight to improve the sensitivity of the multi-modal inversion network model to temperature change characteristics in the early stage of fire development; S36, introduce the whole modal Drop strategy in the training stage, specifically: in each batch training, input the whole image to zero with a probability of 0.35, which is regarded as a camera failure or smoke shielding working condition; when the image is zeroed, the multi-modal inversion network model automatically triggers the "image availability gate signal" to 0, making the image branch weight zero, thereby preventing the misentry of invalid information; S37, implement high-probability modal Drop operation at the physical input end, Drop operation with a probability of 0.7 to zero the whole physical input vector, only retaining non-leakage statistical noise or wind speed information, specifically: in the "none" mode configuration, the physical input is completely shielded, only retaining the zero vector form; when set to "wind_only" mode, only the wind speed component is retained for ventilation influence modeling; in "full" mode, only used for comparative test, not as the main model configuration; S38, sensor random masking enhancement mechanism, specifically: in each batch, randomly select up to 10% of the sensor channels with a probability of 0.15 and zero them at all time steps; at the same time, the corresponding temperature difference channels are also zeroed to keep the input feature dimension consistent; S39, use a consistency distillation mechanism to maintain the consistency of model output results under different modal input conditions to improve the stability and generalization performance of the multi-modal inversion network model under multi-source data missing conditions, specifically: in each training process, the multi-modal inversion network model inputs the same batch of training samples in two forms: one is complete modal input, including temperature sensor sequence, temperature image and physical parameters; the other is shielding modal input, which simultaneously zeros the image and physical parameters, only retaining the sensor modal; wherein the former is the student model and the latter is the teacher model, both of which are calculated in parallel under the same structure to output results; S310, in the training process, various losses are dynamically balanced by uncertainty weighting, including HRR regression, position classification, position regression, physical consistency and consistency distillation; the optimizer uses AdamW algorithm, the initial learning rate is 1x10 -4 , the weight decay coefficient is 1x10 -4 , and the learning rate is automatically adjusted by the ReduceLROnPlateau scheduler based on the validation set loss; during the training process, mixed precision training is used to reduce the memory occupation, and the gradient clipping upper limit is set to 1.0 to prevent gradient explosion; when the total loss of the validation set does not improve for a preset number of rounds, the early stopping mechanism is triggered to avoid overfitting.

9. The tunnel fire inversion method based on multi-modal fusion and physical constraints according to claim 1, characterized in that, Step S4 comprises: S41, input data parsing and feature construction: the multi-modal inversion network model first reads the temperature sensor data CSV file obtained by fire simulation or field collection, which contains time series and multiple temperature measurement point records; the multi-modal inversion network model automatically identifies the time column and the temperature column, and selects equidistant sensor columns for analysis according to the number of sensors during training; after reading, the multi-modal inversion network model performs numerical cleaning on all temperature data and generates a time series vector, then constructs a feature tensor according to the sliding window length W=2 in the training stage: each time consists of the current and previous time temperature values and their time difference to form a feature matrix [W,2S], and the feature mean and standard deviation saved during training are used for normalization to ensure that the input distribution is consistent with the training stage; S42, model loading and inference configuration: the system loads the trained multi-modal fusion inversion network model from the specified model directory, including the weight file with pth suffix and the normalization parameter file with npz suffix; the inference process runs in a frozen state, and the multi-modal fusion inversion network model parameters remain unchanged, only forward calculation is performed; the system supports automatic device selection, and runs in CUDA mode when a GPU is detected, otherwise executes on CPU; during inference, the image and physical modal inputs are disabled, and a zero tensor is automatically input and the gating mechanism is enabled to maintain consistency with the training anti-leakage; the sensor coordinates are automatically generated according to the tunnel length and spacing for fire source position grid definition and probability calculation; S43, real-time per-time inference and result smoothing: the system inputs the feature tensor per time and performs forward calculation to obtain the fire power prediction value and the fire source position prediction value, wherein the fire power prediction value is output by the fire power regression head, and the fire source position prediction value is determined by the combination of the coarse classification and fine regression two-level networks; During the inference process, the system retains the LSTM hidden state between adjacent time points to maintain time continuity; at the same time, in order to suppress high-frequency fluctuations, the system uses the exponential moving average algorithm to smooth the power and position prediction results at consecutive time points, improving the stability and physical consistency of the curve; S44, result file generation and visualization output: the system automatically outputs the following result files after completing the inference: Main result file, CSV format: contains the timestamp, fire power prediction value, fire source position prediction value and original prediction result at each time; Probability matrix file, NPZ format: stores the fire source position probability matrix at each time, tunnel discrete grid coordinates and time series; Wide table format CSV file: each column corresponds to the tunnel position, and each row corresponds to the time point, which is used to directly generate the heat map; Long table format CSV file: contains three columns, time, position and probability, which is used for statistical and drawing analysis; The fire power curve graph takes time as the horizontal axis and power as the vertical axis to show the fire development trend; the fire source position probability heat map takes the tunnel position as the horizontal axis and the probability as the vertical axis to represent the probability of fire source occurrence at each position, and intuitively reflects the space-time migration process of the fire source in the tunnel vertical direction.

Citation Information

Patent Citations

  • Drilling overflow prediction combination model based on deep learning and model timely silent updating and transfer learning method

    CN117171700A

  • Tunnel fire source parameter inversion method and system based on data driving

    CN117787104A