Vehicle condition estimation method and device based on multi-modal data fusion, equipment, storage medium and product
By using a multimodal data fusion method, non-image and image data of vehicles are acquired to generate multimodal fusion features, which solves the shortcomings of single-modal data detection in existing technologies and enables accurate assessment of vehicle conditions and timely detection of anomalies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SOUTH CHINA UNIV OF TECH
- Filing Date
- 2025-07-29
- Publication Date
- 2026-07-24
AI Technical Summary
Existing technologies for vehicle anomaly detection often rely on single-modal data and lack in-depth fusion analysis of multi-source heterogeneous information, resulting in a high false alarm rate and affecting the timeliness of fire response.
A multimodal data fusion method is adopted to acquire non-image data and image data of at least two modalities of the vehicle. Through feature extraction, location-aware feature generation, joint modeling and local-global network fusion, multimodal fusion features are generated to predict vehicle conditions.
It improves the accuracy of vehicle condition detection, enabling timely detection of vehicle anomalies and enhancing the timeliness and accuracy of fire response.
Smart Images

Figure CN121010770B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of vehicle control technology, and more specifically, to a method, apparatus, equipment, storage medium, and product for vehicle condition prediction based on multimodal data fusion. Background Technology
[0002] With the rapid increase in the number of new energy vehicles, the safety risks caused by thermal runaway of power batteries have become a new challenge, especially in relatively enclosed indoor environments where the threat is even more severe. Therefore, real-time assessment of vehicle status is particularly necessary. However, current vehicle anomaly detection technology has significant shortcomings. Although multimodal data assessment models have been introduced, most technologies analyze single-modal data such as temperature and smoke independently, and then simply summarize the assessment results, lacking in-depth fusion analysis of multimodal data. This significantly reduces the accuracy of identifying abnormal vehicle states. Summary of the Invention
[0003] Based on this, the present invention provides a vehicle condition prediction method, device, equipment, storage medium and product based on multimodal data fusion, to solve the problem of insufficient state recognition accuracy caused by independent evaluation of vehicle state by each modality data in the prior art.
[0004] To achieve the above objectives, embodiments of the present invention provide a vehicle condition prediction method based on multimodal data fusion, comprising:
[0005] Acquire real-time monitoring data of the vehicle; wherein the real-time monitoring data includes non-image data and image data in at least two modalities;
[0006] Feature extraction and encoding, as well as dimension unification processing, are performed on the non-image data and the image data respectively to obtain non-image features and image features;
[0007] Spatial location information is embedded into the image features to generate location-aware features;
[0008] The position-aware features of the image data of each modality are jointly modeled based on the attention mechanism and the hybrid expert module to generate intermediate features of the image data of each modality;
[0009] Based on a local-global network, all the intermediate features are fused to generate image fusion features;
[0010] The image fusion features and the non-image features are fused to generate multimodal fusion features;
[0011] The vehicle condition prediction result is predicted based on multimodal fusion features.
[0012] To achieve the above objectives, embodiments of the present invention also provide a vehicle condition prediction device based on multimodal data fusion, comprising:
[0013] The data acquisition module is used to acquire real-time monitoring data of the vehicle; wherein, the real-time monitoring data includes non-image data and image data in at least two modalities;
[0014] The feature extraction module is used to perform feature extraction and encoding as well as dimension unification processing on the non-image data and the image data respectively, to obtain non-image features and image features;
[0015] The location embedding module is used to embed spatial location information into the image features to generate location-aware features;
[0016] The joint modeling module is used to jointly model the position-aware features of the image data of each modality based on the attention mechanism and the hybrid expert module, and generate intermediate features of the image data of each modality.
[0017] An image fusion module is used to fuse all the intermediate features based on a local-global network to generate image fusion features;
[0018] The feature fusion module is used to fuse the image fusion features and the non-image features to generate multimodal fusion features;
[0019] The vehicle condition prediction result is predicted based on multimodal fusion features.
[0020] To achieve the above objectives, embodiments of the present invention also provide a vehicle condition prediction device based on multimodal data fusion, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the vehicle condition prediction method based on multimodal data fusion as described in any of the above embodiments.
[0021] To achieve the above objectives, embodiments of the present invention also provide a computer-readable storage medium, the computer-readable storage medium including a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to execute the vehicle condition prediction method based on multimodal data fusion as described in any of the above embodiments.
[0022] To achieve the above objectives, embodiments of the present invention also provide a computer program product, including a computer program / instructions, which, when executed by a processor, implement the vehicle condition prediction method based on multimodal data fusion as described in any of the above embodiments.
[0023] Compared with existing technologies, the vehicle condition prediction method, apparatus, device, storage medium, and product based on multimodal data fusion disclosed in this invention first acquires non-image data and image data of at least two modalities of the vehicle, and performs feature extraction and encoding, as well as dimensionality unification processing, on the non-image data and the image data respectively to obtain non-image features and image features. Then, spatial location information is embedded into the image features to generate location-aware features, and the location-aware features of the image data of each modality are jointly modeled based on an attention mechanism and a hybrid expert module to generate intermediate features of the image data of each modality. Next, all the intermediate features are fused based on a local-global network to generate image fusion features. The image fusion features and the non-image features are then fused to generate multimodal fusion features. Finally, the vehicle condition prediction result is predicted based on the multimodal fusion features. Therefore, this invention can obtain multimodal fusion features by deeply fusing image data of multiple modalities and then fusing it with non-image data. The multimodal fusion features are then used to accurately assess the vehicle condition, improving the accuracy of vehicle condition detection and enabling timely detection of vehicle anomalies. Attached Figure Description
[0024] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a flowchart illustrating a vehicle condition prediction method based on multimodal data fusion according to an embodiment of the present invention.
[0026] Figure 2 This is a schematic diagram of a multimodal data fusion network structure provided in an embodiment of the present invention;
[0027] Figure 3 This is a schematic diagram of the working process of an intelligent fire extinguishing system for parking lots provided in an embodiment of the present invention;
[0028] Figure 4 This is a schematic diagram of a vehicle condition prediction device based on multimodal data fusion provided in an embodiment of the present invention;
[0029] Figure 5 This is a schematic diagram of the structure of a vehicle condition prediction device based on multimodal data fusion provided in an embodiment of the present invention. Detailed Implementation
[0030] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0031] Currently, there are significant shortcomings in parking lot anomaly detection and handling technologies. Most existing research and applications rely on single-modal data, such as temperature sensors or smoke detectors, lacking in-depth fusion analysis of multi-source heterogeneous information (including visible light and infrared images, smoke concentration change sequences, etc.). This not only leads to a high false alarm rate but also affects the timeliness of fire response.
[0032] Based on this, this embodiment provides a vehicle condition prediction method based on multimodal data fusion, see [link to relevant documentation]. Figure 1 The vehicle condition prediction method based on multimodal data fusion includes steps S1 to S7:
[0033] S1. Acquire real-time monitoring data of the vehicle; wherein, the real-time monitoring data includes non-image data and image data in at least two modalities;
[0034] S2. Perform feature extraction and encoding, as well as dimension unification processing, on the non-image data and the image data respectively to obtain non-image features and image features;
[0035] S3. Embed spatial location information into the image features to generate location-aware features;
[0036] S4. Based on the attention mechanism and the hybrid expert module, the position-aware features of the image data of each modality are jointly modeled to generate intermediate features of the image data of each modality;
[0037] S5. Based on a local-global network, all the intermediate features are fused to generate image fusion features;
[0038] S6. The image fusion feature and the non-image feature are fused to generate a multimodal fusion feature;
[0039] S7. Predict the vehicle condition forecast based on the multimodal fusion features.
[0040] Specifically, it is worth noting that the vehicle condition prediction method based on multimodal data fusion can be executed by the vehicle itself, or it can be applied to an intelligent fire suppression system in a parking lot to fuse multimodal data and achieve vehicle condition assessment. See also Figure 2 The diagram shows a multimodal data fusion network structure. The data fusion process is described in detail below:
[0041] In step S1, the non-image data can be smoke data (Smoke, S) or other time series data used to reflect changes in smoke concentration, and the image data can include infrared images (I) or visible light images (V).
[0042] In step S2, multimodal data feature extraction and encoding are performed. Specifically, for visible light images, infrared images, and smoke data, respectively, the following methods are used: (Right now The input sequence length and feature dimension are represented by ). Given the significant differences in information representation and data structure among the three modalities, a multi-branch encoder was designed. Corresponding encoders were used to extract and encode features from visible light images, infrared images, and ambient smoke, respectively. The features from different modalities were then projected onto a common dimension d, yielding their respective high-dimensional feature representations. As shown in equation (1).
[0043]
[0044] Among them, (1) for visible light images, the encoder is an image encoder, which can be a convolutional neural network (CNN), such as a residual network (ResNet50), an efficient network (EfficientNet), a sliding window transformer (Shifted Window Transformer, Swing Transformer), etc., or a pre-trained model, such as a pre-trained model of the contrastive language-image pretraining (CLIP) model: visual transformer (32x32)(Vision Transformer Base-32, ViT-B-32), visual transformer (16x16)(Vision Transformer Base-16, ViT-B-16), visual transformer (14x14)(Vision Transformer Base-14, ViT-B-14), etc., to improve the generalization ability of features; Transformer is a deep learning model based on attention mechanism. (2) Infrared data reflects thermal distribution information, and its numerical characteristics differ from those of visible light. The encoder for infrared images is an image encoder, which can be a CNN with an attention mechanism (such as a Squeeze-and-Excitation ResNet (SE-ResNet) or a Convolutional Block Attention Module (CBAM)). It focuses on mining the details of the target heat source edge and temperature gradient, and combines multi-scale convolution to achieve fine recognition of hotspot areas. (3) For the sequence data of on-site smoke (i.e., smoke data), which contains multi-dimensional information such as gas concentration changes and current signals, the encoder is a sequence encoder. The specific types can be one-dimensional convolutional neural networks (1D-CNN), recurrent neural networks (RNN), long short-term memory networks (LSTM), gated recurrent units (GRU), or Transformers, etc., to capture dynamic changes in the time dimension. In addition, multi-scale convolution and attention mechanisms can also be combined to enhance feature representation capabilities.
[0045] In step S3, positional encoding is performed. Specifically, to preserve spatial structure information in the image, absolute positional encoding based on sine and cosine functions is introduced to encode image features. (Image features of visible light images) Image features of infrared images Spatial location information enhancement is performed to effectively improve the downstream module's ability to perceive changes in spatial layout. Position encoding is added to the high-dimensional features obtained after encoding visible light and infrared images to preserve the image's positional information in spatial structure, as shown in Equation (2).
[0046]
[0047] In the formula, This represents the location-aware features obtained by adding location coding to the image features of visible light images and infrared images; This represents the location encoding function, which generates a fixed embedding vector for each spatial location.
[0048] The positional encoding values are generated by alternating sine and cosine functions, with their frequency determined by the dimension index j. sy Decision (j) sy ∈[0,d / 2-1]). Specifically, for each position i in this sequence of image features... wz (i wz (∈[0,T-1], where T is the sequence length) By combining sine and cosine functions of different frequencies, a unique encoding vector is generated, thereby achieving effective representation of spatial location information. The specific calculation method of location encoding is as follows:
[0049]
[0050] In steps S4 to S6, an n-layer Cross-Modality Expert Transformer (CMET) is used to combine the position-aware features of the visible light image and the infrared image. and These are respectively designated as Mode 1 (Mode u1) and Mode 2 (Mode u2) in the CMET module, and deep processing is performed on them. The advantages of Multi-Head Attention (MHA) and Mixture of Experts (MOE) are cleverly combined to achieve joint modeling of the encoded features of the two modes, outputting... and These are used as intermediate features for the visible light image and the infrared image, respectively. The Adaptive Selection and Fusion (ASF) module uses the two outputs obtained from the CMET module to perform intermediate feature processing. and (The intermediate features of the visible light image and the intermediate features of the infrared image) are further processed with adaptive selection to obtain image fusion features. Finally, the image features are fused based on the Dynamic Adaptive Fusion (DAF) module. By fusing non-image features, multimodal fusion features are generated.
[0051] In step S7, vehicle condition assessment and vehicle spontaneous combustion time estimation are performed. Specifically, the final fused feature output of the DAF module... Dimensional mapping is performed through a fully connected (FC) layer to generate prediction results (i.e., vehicle condition estimation results) for assessing the condition of vehicles in existence. The multimodal data fusion network, by fusing multimodal data, extracting temporal features, and combining importance weighting between modes, can effectively capture deep dynamic changes in vehicle status, achieving accurate predictions of key indicators such as whether a vehicle is in normal operating condition and whether it is approaching the combustion stage, providing strong support for subsequent scheduling management and safety early warning.
[0052] For example, by using the above-mentioned multimodal data fusion method to assess the vehicle condition, the vehicle condition of vehicles with spontaneous combustion risk can be classified into three states: the early stage of thermal runaway, the thermal runaway stage, and the vehicle combustion stage.
[0053] Compared with existing technologies, the embodiments of the present invention can obtain multimodal fusion features by deeply fusing image data of multiple modalities and then fusing it with non-image data. The multimodal fusion features are used to accurately assess the vehicle status, improve the accuracy of vehicle condition detection, and enable timely detection of vehicle anomalies.
[0054] In a preferred embodiment, the step of fusing all the intermediate features based on a local-global network to generate image fusion features includes:
[0055] By concatenating all the intermediate features, a combined feature is obtained;
[0056] The combined features are processed using a local-global network to extract multi-scale features;
[0057] Image fusion features are generated based on the multi-scale features and all the intermediate features.
[0058] Specifically, see Figure 2 This document describes the data processing flow of the ASF module. The ASF module dynamically evaluates and selects information from different modalities for optimal feature fusion. Leveraging the context of the specific task, this module automatically adjusts the weights of each modality, ensuring that the most relevant and useful features are prioritized, thereby effectively reducing redundant information and noise. In this way, the model not only better captures key information but also improves overall prediction performance and robustness, achieving more accurate results. This fusion mechanism enables the network to flexibly handle diverse data inputs, enhancing the model's adaptability to complex scenarios. Specifically, the ASF module is used to process the two outputs obtained from the CMET module. and (The intermediate features of the visible light image and the intermediate features of the infrared image) are further processed with adaptive selection to obtain image fusion features. To achieve the selective fusion of the two modes, the fusion output is shown in equation (5).
[0059]
[0060] In the formula, LGNet(·) represents Local-Global Network (LGNet), which achieves efficient extraction and fusion of multi-scale features by combining local feature processing and global modeling. It can effectively capture fine-grained local details and long-range global context. Its calculation formula in general form is shown in Equation (6); Concat(·) represents concatenation, which connects and merges multiple tensors along a specified dimension.
[0061] LGNet(x)=σ(Local(x)+Global(x)) (6)
[0062] Local(x) = LN(FFN(LN(x))) (7)
[0063] Global(x)=MHA(LN(FFN(LN(x))) / +AAP(MHA(LN(FFN(LN(x))))) (8)
[0064] In the formula, σ(·) represents the Sigmoid function; MHA(·) represents the multi-head attention mechanism, and its calculation formula in the general form is shown in equation (9); FFN(·) consists of FC, GeLU, and FC. Figure 2In this context, LN&FC means LN followed by FC, and FC&LN means FC followed by LN. GeLU represents the activation function. LN represents layer normalization, and FC represents a fully connected layer. AAP(·) represents Adaptive Average Pooling (AAP), which further optimizes training stability and computational efficiency, making the model more robust to complex data. The general formula for calculating the MHA module is as follows:
[0065]
[0066] Among them, Q x K represents the query vector in its general form. x V represents the key vector in its general form. x d represents a value vector in its general form. k Represents the key dimension in its general form; the Softmax function is a soft maximization function.
[0067] In a preferred embodiment, fusing the image fusion features and the non-image features to generate multimodal fusion features includes:
[0068] The image fusion features and the non-image features are concatenated to generate concatenated features;
[0069] The contextual temporal information of the spliced features is extracted, and image modality weights and non-image modality weights are learned based on the contextual temporal information. The image fusion features and the non-image features are then weighted and summed based on the image modality weights and the non-image modality weights to generate weighted fusion features.
[0070] The splicing features are subjected to a nonlinear transformation to generate a residual vector;
[0071] Multimodal fusion features are generated based on the contextual temporal information, the weighted fusion features, and the residual vector.
[0072] Specifically, see Figure 2 The specific implementation method for the DAF module fusion step is as follows:
[0073] The DAF module can efficiently process the output of the ASF module. High-dimensional feature representation of smoke data The integration between them, Belongs to spatial visual information, Belonging to time series features, this module integrates three mechanisms: time series modeling, dynamic allocation of modal importance, residual connection and nonlinear adaptation. (1) Time series modeling: Using bidirectional gated recurrent units (Bi-GRU), splicing features are extracted. (1) Contextual temporal information to improve the ability to perceive dynamic changes; (2) Dynamic weighted fusion: learn image modality weights and non-image modality weights directly from the Bi-GRU output through a fully connected network (it is worth noting that the Softmax normalization technique is used to process the weights in this process), realize adaptive modeling of modality importance, obtain dynamic weight, and ensure that the contribution ratio of each modality is dynamically adjusted in different scenarios; (3) Residual connection and nonlinear adaptation: use a residual connection branch (including a fully connected layer (FC), LN and GELU activation function) to splice features Performing linear transformations and nonlinear activations preserves the original information while adapting it to the target space, which helps stabilize training and enhance representation capabilities.
[0074] Output of the DAF module The calculation method is shown in equation (10).
[0075]
[0076] In the formula, Indicates multimodal fusion features, This represents the output result (i.e., contextual timing information) obtained through Bi-GRU fusion; This represents the output result after dynamic modality weighted fusion (i.e., weighted fusion features); This represents the residual vector obtained by performing a nonlinear transformation on the concatenated input. Finally, it is processed through a fully connected (FC) layer. Dimensional mapping is performed to generate prediction results, which are used to assess the vehicle's condition and the time of spontaneous combustion.
[0077] (1) Temporal modeling: the spliced sequence The input is fed into a bidirectional gated loop unit (Bi-GRU) to obtain the forward output. and reverse output Summing them together yields the output result. The formula is shown below.
[0078]
[0079] (2) Dynamic weighted fusion
[0080] Learn weight vectors using a perceptron network
[0081]
[0082] in, Represents image modal weights. Represents non-image modality weights;
[0083] Merge based on weights:
[0084]
[0085] (3) Residual connection and nonlinear adaptation
[0086] The residual vector is obtained by performing a nonlinear transformation on the spliced features:
[0087]
[0088] Among them, W r The weights, b, represent the nonlinear transformation weights. r This represents the bias of the nonlinear transformation.
[0089] In a preferred embodiment, the joint modeling of the position-aware features of the image data for each modality based on the attention mechanism and the hybrid expert module to generate intermediate features of the image data for each modality includes:
[0090] For the image data of the first modality, a first query vector is generated based on the first location-aware feature, and a first key vector and a first value vector are generated based on the second location-aware feature; wherein, the first location-aware feature is the location-aware data of the image data of the first modality, and the image data of the first modality is one of the image data of all modalities;
[0091] Based on the attention mechanism, a first multi-attention output is generated according to the first query vector, the first key vector, and the first value vector;
[0092] After performing a residual connection between the first multi-attention output and the first position-aware feature, a layer normalization operation is performed to obtain the first layer normalized output.
[0093] The first layer normalized output is input into the hybrid expert module for processing to obtain the first hybrid expert output.
[0094] After performing residual processing on the first hybrid expert output and the first layer normalized output, a layer normalization operation is performed to obtain the first intermediate feature.
[0095] Specifically, see Figure 2In this stage, the CMET module utilizes an n-layer Cross-Modality Expert Transformer (CMET) to integrate the position-aware features of visible light and infrared images. and These are respectively used as Mode 1 (Mode u1) and Mode 2 (Mode u2) in the CMET module, and deep processing is performed on them. The core of this module lies in the ingenious integration of the advantages of multi-head attention mechanism and hybrid expert module to achieve joint modeling of the two modality coding features, outputting Out 1 and Out 2, which serve as intermediate features of visible light image and infrared image, respectively.
[0096] Specifically, the CMET module first utilizes a cross-modal attention mechanism to achieve information exchange and semantic alignment between visible light and infrared images. Then, the MOE module is introduced to dynamically model the fused features. The MOE module includes a router and a gating unit. The router selectively activates multiple closely related expert subnetworks based on the input data; the gating unit assigns weights to each expert and integrates their outputs. In this way, the MOE module not only significantly enhances the model's nonlinear expressiveness and generalization performance but also possesses excellent scalability and computational efficiency, dynamically adjusting the processing path based on the input and effectively avoiding resource waste. Furthermore, the overall structure ensures the stability of deep network training and the continuity of information flow through residual connections and layer normalization (LN). Residual connections are a structure that directly adds the input of one layer in the network to the output of a subsequent layer. The core idea of this connection method is to allow signals in the network to bypass certain layers and propagate directly, thereby alleviating the problems of gradient vanishing or gradient exploding in deep network training, and also helping the network learn more complex features.
[0097] Taking the transfer of infrared image (I) information to visible light image (V) as an example, denoted as "I→V". In the Multi-HeadAttention module, all dimensions are unified to maintain the consistency of the feature space. The specific data processing procedure of the CMET module is shown in the following equation:
[0098]
[0099]
[0100]
[0101] In the formula, It is a position-aware feature of visible light images, As This represents the cross-modal attention mapping from the infrared image mode to the visible light image mode in the i-th layer, calculated by the Multi-Head Attention module. The detailed calculation principle is shown in Equation (19). MOE [i] Let L represent the calculation formula for the i-th layer MOE module. The detailed calculation principle is shown in equation (20). LN represents layer normalization.
[0102]
[0103] In the formula, X α X represents the input vector or feature matrix of mode α. β d represents the input vector or feature matrix of mode β. k This represents the scaling factor for the dimensions of the key vector K and the query vector Q, used to adjust the numerical range of the dot product. Figure 2 In this context, V′ represents the value vector; Q α The query vector representing mode α. K β The key vector representing mode β. V β The vector representing the value of mode β. in, and These are the query weight matrix, key weight matrix, and value weight matrix, respectively. Where, d α d represents the feature dimension of mode α. β d represents the feature dimension of the modality β. k d represents the feature dimension of the key vector. v The characteristic dimension of a value vector.
[0104]
[0105] In the formula, The input vector at time step t represents the i-th layer, which is the intermediate hidden state obtained after passing through the MHA module and layer normalization; This represents the output vector of the i-th layer at time step t (i.e., the output of the MOE module); FFN j (·) represents the feedforward neural network of the j-th expert, used to perform nonlinear transformations on the input. The specific formula can be expressed as FFN(x)=FC(GeLU(FC(x)))=GeLU(xW1+b1)W2+b2, where GeLU is the activation function, W1 is the first-layer weight matrix, b1 is the first-layer bias term, W2 is the second-layer weight matrix, and b2 is the second-layer bias term; mN represents the total number of experts in the system, where m is the number of expert groups, and each group contains N experts; K sThis represents the number of shared experts who are active at all time steps; g j,t This represents the gating weight of the j-th expert at time t, used to control the activation state of that expert. It's worth noting that, for simplicity, FFN is commonly used to represent the components of FC, GeLU, and FC, while FC represents a fully connected layer.
[0106] Formula (20) consists of two parts: (1) Shared expert part: the first K s (2) Dynamic selection part: from the remaining mN-K experts are forcibly activated, and their outputs are directly added together to ensure that the model always retains a certain shared feature extraction capability. s Among the experts, based on the gating weight g j,t A subset of experts is dynamically selected to participate in the calculation, and their outputs are integrated through a weighted summation.
[0107] Gating weight g j,t The calculation formula is as follows:
[0108]
[0109]
[0110] In the formula, s j,t The value represents the activation score of the j-th expert at time step t, obtained after Softmax normalization; Topk(·,mK-K) s () indicates selecting the top mK-K scores from the candidate experts. s The number of experts; mK represents the total number of experts allowed to be activated at each time step (including shared experts), where K is the number of experts activated in each group; This represents the original calculation basis for the expert activation score, expressed through the input vector. With expert gating vectors The inner product is implemented. The gating vector represents the j-th expert in the i-th layer, which is learned and used to measure the degree of matching between the input and the expert. The Softmax operation ensures that the sum of the activation scores of all candidate experts is 1, thereby normalizing the weights.
[0111] In a preferred embodiment, the non-image data is on-site smoke information, and the image data includes infrared images and visible light images.
[0112] In a preferred embodiment, the method further includes:
[0113] If the vehicle condition assessment result includes the time of vehicle spontaneous combustion, then the vehicle is the target vehicle;
[0114] The system obtains the current vehicle location of the target vehicle in the parking lot, the location of the safe zone in the parking lot, the road layout information of the parking lot, and the information of key equipment on both sides of the road.
[0115] Starting from the current vehicle location and ending at the location of the safe zone, path planning is performed based on the road layout information to obtain at least one candidate path and its estimated travel time.
[0116] The estimated travel time is limited to be less than the time of vehicle spontaneous combustion, and a target route is selected from the candidate routes;
[0117] The target vehicle is moved based on the target path.
[0118] It is worth noting that when assessing vehicle condition based on multimodal fusion features, the time of spontaneous combustion of vehicles with a risk of spontaneous combustion is also predicted. The time of spontaneous combustion refers to the duration between the current moment and the estimated moment when open flames may appear.
[0119] For example, cameras can be installed in parking lots to collect parking lot information, thereby determining the location of all vehicles, the location of all safe zones, the road layout information of the parking lot, and the information of key equipment on both sides of the road. Key equipment includes critical facilities and vehicles, including but not limited to electrical boxes, backup generators, and charging piles. The fire risk of key equipment is not limited to damage to the equipment itself, but may also amplify safety accidents by igniting flammable materials, damaging fire-fighting facilities, blocking escape routes, or providing conditions conducive to combustion. Therefore, these devices require special attention. Key equipment information includes, but is not limited to, the type of key equipment, the location information of key equipment, and the distance of key equipment from the center of the road. Optionally, the vehicle condition prediction method based on multimodal data fusion can be applied to the intelligent fire extinguishing system of the parking lot. The vehicle location can also be reported by the vehicle to the system, and the road layout information of the parking lot can be determined based on the architectural design drawings of the parking lot.
[0120] For example, the real-time monitoring data in step S1 is obtained from sensing devices, which may include sensing devices pre-deployed in the parking lot, sensing devices deployed on vehicles, and / or sensing devices deployed on car-moving robots. Sensing devices may include cameras, smoke detectors, and temperature sensors, etc., and the real-time monitoring data may include image data (such as visible light images, infrared images, etc.) and non-image data (such as smoke data, temperature data, etc.). The sensing devices are not limited to the specific types mentioned above.
[0121] For example, when the temperature sensor for the vehicle's power battery detects a temperature exceeding a preset threshold, or when a smoke detector in a parking lot detects smoke coming from the vehicle, the vehicle is considered to have a risk of spontaneous combustion. It's worth noting that the method for determining spontaneous combustion risk is not limited to the specific examples mentioned above; other methods can be used depending on the actual situation. When a risk of spontaneous combustion is determined, the vehicle's current state is analyzed based on real-time monitoring data, and the estimated time of spontaneous combustion is assessed.
[0122] It is worth noting that a safe zone refers to an area with a certain degree of isolation. This area can be pre-arranged or planned based on the current parking situation and the placement of critical equipment. This area can isolate the target vehicle and critical equipment, reducing the possibility of fire spreading to other equipment. Road layout information records the location, direction, length, and width of each road, while critical equipment information records the location of critical equipment and / or its distance from the center of the road. Therefore, the distance between the center of each candidate path and the critical equipment on both sides of the road can be determined based on the road layout information and the critical equipment information.
[0123] Specifically, the vehicle can be either a new energy vehicle or a gasoline-powered vehicle. Since a fire in a vehicle can easily spread to other equipment (such as other new energy vehicles), potentially causing a cluster fire, it is necessary to move the vehicle to a safe area before it catches fire to isolate it and reduce the possibility of the fire spreading. Furthermore, unexpected events may occur during the vehicle relocation process, such as premature ignition. To deal with such sudden situations, it is necessary to choose an optimal route to move the vehicle, for example, choosing a route that takes critical equipment on both sides of the road as far away from the center of the road as possible to reduce the risk of critical equipment being ignited.
[0124] Optionally, a subset of paths can be selected from all candidate paths. The selected paths must meet the following conditions: the estimated travel time is less than the time of vehicle fire, and the shortest distance from the critical equipment to the center of the road is greater than a set distance threshold. Then, from the selected paths, the path with the largest average distance from all critical equipment to the center of the road is chosen as the target path. Alternatively, paths with an estimated travel time less than the time of vehicle fire can be selected from all candidate paths. The shortest distance from the critical equipment to the center of the road is determined for each selected candidate path, and the path with the largest shortest distance is chosen as the target path. It is worth noting that the selection method for the target path is not limited to the specific methods described above and can be chosen according to the actual situation. The target vehicle can either start automatically and move to the corresponding safe area according to the target path, or a vehicle-moving robot can move the target vehicle to the corresponding safe area to isolate the target vehicle.
[0125] For example, taking the vehicle condition prediction method based on multimodal data fusion applied to a parking lot intelligent fire suppression system, the parking lot intelligent fire suppression system can be divided into a vehicle anomaly monitor, an equipment management platform, and a vehicle relocation robot. These modules cooperate with each other to form a complete emergency response closed-loop system. The main functions of each module of the parking lot intelligent fire suppression system are as follows:
[0126] Vehicle anomaly monitors: Multiple vehicle anomaly monitors are deployed in the parking lot, each covering a small area to achieve blind spot coverage. Each vehicle anomaly monitor can integrate a visible light camera, an infrared thermal imaging camera, a smoke detector, and a wireless signal receiver, enabling it to accurately identify the status of each parking space and obtain the precise location coordinates of abnormal vehicles. The vehicle anomaly monitors have the following functions: (1) Real-time monitoring and synchronization of information from visible light cameras, infrared thermal imaging cameras, and smoke detectors in the parking lot to the equipment management platform; (2) Receiving pre-ignition or abnormal signals from vehicles (such as new energy vehicles or fuel vehicles) and transmitting the signals to the equipment management platform.
[0127] It's worth noting that the visible light camera can capture high-resolution color images, identifying vehicle appearance and behavioral characteristics for abnormal behavior recognition; the infrared thermal imaging camera monitors the temperature distribution of the vehicle and its environment, identifying overheating, short circuits, and other thermal anomalies in real time; and the smoke detector detects tiny smoke particles in the air, providing early warning of fires in the initial stages of combustion. The visible light camera, infrared thermal imaging camera, and smoke detector work together to improve monitoring accuracy and response speed. The vehicle anomaly monitor has a built-in Bluetooth module, WiFi receiver module, and local broadcast system (such as Zigbee) for efficient, low-latency short-range data reception. Considering the actual communication conditions on-site, the device management platform supports access to Low-Power Wide-Area Network (LPWAN) technologies, such as Long Range Radio (LoRa) and Narrow Band Internet of Things (NB-IoT), to enhance network coverage and communication redundancy. In environments with stable cellular networks, the vehicle anomaly monitor and the vehicle can also upload data in real time via 4G / 5G networks, forming a dual guarantee of local detection and remote management, significantly improving system robustness and response speed. Furthermore, the communication between the vehicle anomaly monitor and the device management platform adopts a multi-mode solution, flexibly selected according to the actual application scenario, including Ethernet (preferred), WiFi, LoRa, NB-IoT, and 4G / 5G cellular networks. Ethernet and WiFi are suitable for scenarios with high bandwidth requirements and real-time transmission of images and videos, while LoRa and NB-IoT are suitable for low-power, long-distance status signal transmission, and cellular communication is suitable for compensating for network coverage blind spots and remote connections. The parallel deployment of multiple communication channels constitutes a redundancy mechanism, greatly improving the stability and availability of the system and ensuring timely uploading and rapid response of anomaly data.
[0128] Equipment Management Platform: The equipment management platform can be deployed on a local or cloud server, with the operating terminal located in the fire control room. It has the following functions: (1) It can receive on-site images (visible light images and infrared images) and smoke information transmitted by the vehicle anomaly monitor in real time via wired or wireless means, and supports remote observation function; (2) It can receive wireless signals from vehicles (such as new energy vehicles or fuel vehicles); (3) If a pre-ignition or abnormal signal is received from the vehicle, the equipment management platform will simultaneously use the vehicle anomaly monitor to reconfirm the location coordinates of the abnormal vehicle; (4) If no pre-ignition or abnormal signal is received from the vehicle, but the vehicle anomaly monitor identifies an abnormal situation (such as flame light, combustion smoke, or the infrared thermal imaging camera detects that the vehicle temperature exceeds the preset temperature threshold), the equipment management platform determines the abnormal vehicle. (5) Send alarm reminders to the on-duty personnel so that they can confirm and handle the alarm; (6) Send a buzzer alarm signal to the alarm in the abnormal area; (7) In the early stage of the warning, start the moving robot and provide it with the location coordinates of the abnormal vehicle; (8) Receive multimodal data such as visible light images, infrared images, and on-site smoke collected by the moving robot in real time, and at the same time receive the abnormal vehicle status assessment results, the estimated time of vehicle spontaneous combustion and the time to move to the safe area (i.e., the estimated passage time); (9) On-duty personnel can directly control the moving robot through the equipment management platform in any scenario and at any time; (10) Plan the path based on the estimated passage time and the time of vehicle spontaneous combustion.
[0129] Car Moving Robot: Permanently stationed in the best safe area or vacant parking space, made of heat-insulating coating and high-temperature resistant materials, equipped with multiple sensors (such as visible light camera, infrared thermal imaging camera, smoke detector and lidar, etc.), and also equipped with multi-mode wireless signal receiving and transmitting modules. The car moving robot can be equipped with wheeled or tracked chassis to adapt to different site conditions and heat resistance requirements. Its specific functions are as follows: (1) Path planning. After receiving the start command and abnormal vehicle location coordinates from the equipment management platform, the car moving robot performs global path planning, generates an initial path, moves towards the target vehicle along the initial path, and at the same time uses lidar to scan the environment to achieve dynamic obstacle avoidance and local path replanning. After dynamically bypassing obstacles, it returns to the original planned path until it accurately arrives at the target vehicle. (2) Vehicle condition assessment. After arriving at the target vehicle, based on multimodal data fusion technology, it performs in-depth analysis of the visible light image, infrared image and smoke data of the target vehicle, and performs path planning to obtain multiple candidate paths, the estimated travel time of each candidate path, the estimated time of vehicle spontaneous combustion, etc. (3) Move the target vehicle based on the final selected target path.
[0130] Specifically, see Figure 3 The diagram illustrates the workflow of an intelligent fire suppression system for parking lots. A brief overview of this workflow follows:
[0131] 1. Vehicles (such as new energy vehicles or fuel vehicles) send real-time monitoring data, such as pre-ignition or abnormal signals, to the vehicle anomaly monitor or the equipment management platform in the fire control room via wireless communication. The vehicle anomaly monitor then transmits the real-time monitoring data to the equipment management platform.
[0132] 2. The equipment management platform comprehensively analyzes the received real-time monitoring data. If an abnormal vehicle is detected, its location (i.e., the current position of the abnormal vehicle) is immediately pinpointed and reported to the vehicle relocation robot. Simultaneously, an alarm is sent to the on-duty personnel and the alarm in the area where the abnormal vehicle is located, causing the alarm to sound an alarm. On-duty personnel can confirm and handle the alarm through the real-time video provided by the vehicle anomaly monitor. If the on-duty personnel determine that the alarm was falsely triggered, they will deactivate it. After confirming the alarm, the on-duty personnel can control the vehicle relocation robot through the equipment management platform. It is worth noting that abnormal vehicles refer to vehicles with a risk of spontaneous combustion.
[0133] 3. After receiving the location information of an abnormal vehicle, the moving robot generates the optimal driving route based on a path planning algorithm and heads to the target vehicle's coordinates. Upon arrival, the sensors on the moving robot collect multimodal data, including visible light images, infrared images, and smoke information (such as the on-site smoke concentration), to comprehensively assess the target vehicle's condition. It then automatically enters under the vehicle to align with and lift it to the relocation point. Simultaneously, it uses the path planning algorithm to calculate the distance and estimated travel time to various safe areas, planning multiple candidate routes. The robot then feeds back the vehicle condition assessment results, candidate routes, the time of the vehicle's spontaneous combustion, and the estimated travel time to each safe area to the equipment management platform in real time. This assists on-duty personnel in timely and accurate assessment and decision-making regarding the abnormal vehicle's condition. It is worth noting that the path planning function can also be performed by the equipment management platform.
[0134] 4. The equipment management platform selects the candidate routes from all candidate routes whose estimated travel time is less than the time when the vehicle spontaneously combusts. It further filters the target routes with the goal of maximizing the distance between the center of the road and the key equipment on both sides of the road. Finally, the selected target routes are sent to the vehicle relocation robot, which moves the target vehicle to the corresponding safe area according to the target route. More specifically, the safety zone is divided into optimal safety zones and general safety zones. The general safety zones are further divided into Level 1, Level 2, Level 3, and Level 4 safety zones, in descending order of their safety levels. When the time T1 of the vehicle fire is greater than the estimated travel time T2 of the candidate path corresponding to the optimal safety zone, the target vehicle is moved to the optimal safety zone according to the candidate path corresponding to the optimal safety zone. When the time T1 of the vehicle fire is less than or equal to the estimated travel time T2 of the candidate path corresponding to the optimal safety zone, the highest-level path with an estimated travel time less than the time of the vehicle fire is selected from all candidate paths corresponding to the general safety zones. If multiple paths are selected, the target path is further selected with the goal of maximizing the distance between the center of the road and the critical equipment on both sides of the road, and the target vehicle is moved according to the target path.
[0135] 5. After the vehicle is placed, the vehicle moving robot returns to its standby position, and the subsequent handling is carried out by the firefighters who arrive at the scene.
[0136] Understandably, the architecture of a parking lot intelligent fire suppression system can be composed of a perception layer, a platform management layer, a network service layer, a response layer, and a security application layer. The perception layer primarily involves vehicle anomaly monitors, responsible for detecting vehicle anomalies, and includes visible light cameras, infrared thermal imaging cameras, smoke detectors, and wireless signal receivers. The platform management layer mainly involves the device management platform. The network service layer is responsible for communication between multiple modules, supporting various connection methods such as Ethernet, WiFi, cellular communication (4G / 5G), local broadcast systems, and low-power wide area network technologies (LoRa, NB-IoT). The response layer covers the specific management of safe areas (such as optimal safety areas and general safety areas). The security application layer implements functions such as real-time monitoring, anomaly alarm and location, automatic response, emergency command, evacuation and rescue, and data analysis.
[0137] Compared with existing technologies, when a vehicle is determined to have a risk of spontaneous combustion, the embodiments of the present invention estimate the time of spontaneous combustion, plan multiple candidate routes to a safe area and determine the estimated travel time for each candidate route, and then, with the goal of maximizing the distance between the center of the road and key equipment on both sides of the road, find a target route whose estimated travel time does not exceed the time of spontaneous combustion. Finally, control the vehicle with a risk of spontaneous combustion to travel along the target route to a safe area, achieving risk isolation before the vehicle spontaneously combusts, reducing the possibility of the fire spreading to other vehicles and equipment after the vehicle spontaneously combusts, thereby reducing the risk of a group fire.
[0138] In a preferred embodiment, the security area includes an optimal security area and several general security areas;
[0139] The step of limiting the estimated travel time to be less than the time of vehicle spontaneous combustion, and selecting a target route from the candidate routes, includes:
[0140] If the estimated travel time of the optimal candidate route is less than the time of the vehicle fire, the optimal candidate route shall be the target route; wherein, the optimal candidate route is the candidate route corresponding to the optimal safe zone;
[0141] If the estimated travel time of the optimal candidate route is greater than or equal to the time of vehicle fire, the estimated travel time is limited to be less than the time of vehicle fire. The target route is selected from all general candidate routes with the goal of maximizing the distance between the center of the road and the critical equipment on both sides of the road. The general candidate route is the candidate route corresponding to the general safety zone.
[0142] Specifically, the safe zone includes the best safe zone and several general safe zones.
[0143] It's worth noting that the "optimal safety zone" in a parking lot is designed to address the emergency transfer and isolation of vehicles at risk of spontaneous combustion, ensuring the safety of personnel, property, and the overall environment. The core function of the optimal safety zone is "isolation and control," therefore, it is equipped with a dedicated isolation compartment to physically isolate the vehicle. This compartment is constructed of high-temperature resistant, fireproof, and corrosion-resistant materials, capable of withstanding potential combustion, high temperatures, or explosive impacts from the vehicle for a short period, preventing the fire from spreading to surrounding areas. The optimal safety zone is a temporary parking area for high-risk vehicles and a crucial node for accident prevention and emergency response. It is typically located away from densely populated areas, such as corners of the parking lot or well-ventilated areas, to avoid interfering with the passage of other vehicles and personnel, while also facilitating smoke exhaust and reducing the risk of heat radiation spread. Although this area itself may not have complete automatic fire suppression systems, its location is prioritized for proximity to fire lanes or fire hydrants, ensuring that firefighters can quickly enter and begin firefighting operations if the fire escalates. In addition, the optimal safety zone is equipped with intelligent monitoring equipment such as video surveillance, temperature and humidity sensors, and smoke detectors to achieve remote monitoring and risk prediction in an unmanned state. When a vehicle triggers a pre-ignition alarm in the underground parking lot, the car-moving robot will prioritize moving the vehicle to this area. Furthermore, considering the inherent reaction and movement delays during task execution, and the risk of not being able to promptly move a vehicle to the optimal safety zone after determining a fire risk, the parking lot's intelligent emergency management system introduces a general safety zone to improve the system's response efficiency and flexibility in emergencies. This general safety zone can be set by the equipment management platform based on the current parking situation and fire emergency response capabilities of the parking lot.
[0144] For example, the target path can be determined in one of the following ways: 1. The method is applied to a parking lot intelligent fire extinguishing system. The car moving robot takes the current vehicle position as the starting point and the position of the best safety zone as the ending point, plans the best candidate path and determines the estimated travel time of the best candidate path. If the estimated travel time of the best candidate path is less than the time when the vehicle spontaneously combusts, the car moving robot moves the target vehicle to the best safety zone according to the best candidate path. If the estimated travel time of the best candidate path is greater than or equal to the time when the vehicle spontaneously combusts, the equipment management platform or the car moving robot takes the current vehicle position as the starting point and the positions of each general safety zone as the ending point, plans multiple general candidate paths and determines the estimated travel time of each general candidate path. The estimated travel time is limited to be less than the time when the vehicle spontaneously combusts. The target path is selected from all general candidate paths with the goal of maximizing the distance between the center of the road and the key equipment on both sides of the road. 2. The method is applied to an intelligent fire extinguishing system in a parking lot. The equipment management platform or the car-moving robot takes the current vehicle position as the starting point and the positions of each safety zone as the ending point, plans the optimal candidate path and general candidate paths, and determines the estimated travel time for each candidate path. If the estimated travel time of the optimal candidate path does not exceed the time of vehicle spontaneous combustion, the car-moving robot moves the target vehicle to the optimal safety zone according to the optimal candidate path. If the estimated travel time of the optimal candidate path is greater than or equal to the time of vehicle spontaneous combustion, the estimated travel time is limited to be less than the time of vehicle spontaneous combustion. The target path is selected from all general candidate paths with the goal of maximizing the distance between the center of the road and the key equipment on both sides of the road. 3. The method is applied to the target vehicle. Starting from its current position and ending at the optimal safety zone, the target vehicle plans an optimal candidate path and determines its estimated travel time. If the estimated travel time is less than the time of vehicle combustion, the target vehicle moves to the optimal safety zone along the optimal candidate path, or the target vehicle sends the optimal candidate path to a moving robot, which then moves the target vehicle to the optimal safety zone along the optimal candidate path. If the estimated travel time of the optimal candidate path is greater than or equal to the time of vehicle combustion, multiple general candidate paths are planned starting from the current vehicle position and ending at various general safety zones. The estimated travel time for each general candidate path is determined, ensuring it is less than the time of vehicle combustion. The objective is to maximize the distance between the center of the road and critical equipment on both sides of the road. The target path is then selected from all general candidate paths. It is worth noting that the specific method for determining the target path is not limited to the above example; the execution entity and steps can be set according to the actual situation, and are not limited here.
[0145] Optionally, if the parking lot has an optimal safety zone and general safety zones, path planning can be prioritized based on the optimal safety zone. If the planned path does not meet the requirements, path planning can then be performed on the general safety zones and the target path can be selected, which can reduce the amount of computation to some extent. Alternatively, path planning can be performed on all safety zones simultaneously, or on the optimal safety zone and several general safety zones that are close to the target vehicle simultaneously. This can quickly find a general safety zone that meets the requirements when the optimal safety zone does not meet the requirements.
[0146] Optionally, and it's worth noting, if all safe areas are unreachable, or there is insufficient time to complete the relocation operation (i.e., the estimated travel time for all candidate routes is greater than or equal to the time of the vehicle's spontaneous combustion), the relocation robot's task with the target vehicle will be terminated, and the relocation work will be considered complete. The relocation robot will then smoothly place the target vehicle at its current location and immediately evacuate the scene, awaiting further intervention from firefighters. By default, the relocation robot performs only one relocation task. Regardless of whether the vehicle is successfully moved to a safe area, it will not actively repeat the relocation or intervene in firefighting or other operations. If a new usable safe area subsequently becomes available due to firefighting or vehicle evacuation, the system can reactivate the relocation robot via remote manual command for a supplementary relocation, but this will not affect its default single-task limitation mechanism.
[0147] Furthermore, the optimal safety zone is a dedicated isolation and disposal area pre-arranged in the parking lot; the general safety zone includes several general safety zones of different levels, with higher-level general safety zones having a greater fire spread difficulty than lower-level general safety zones, and higher-level general safety zones having a better fire emergency response capability than lower-level general safety zones.
[0148] If the estimated travel time of the optimal candidate route is greater than or equal to the time of vehicle fire, the estimated travel time is limited to be less than the time of vehicle fire. The target route is selected from all general candidate routes with the objective of maximizing the distance from the center of the road to critical equipment on both sides of the road. This includes:
[0149] If the estimated travel time of the optimal candidate route is greater than or equal to the time of the vehicle fire, select an optional safe area from all the general safe areas whose estimated travel time is less than the time of the vehicle fire.
[0150] The region with the highest security level is selected from the available security regions as a candidate security region;
[0151] With the goal of maximizing the distance between the center of the road and the critical equipment on both sides of the road, the target path is selected from all candidate paths; wherein, the candidate path is the path to be selected corresponding to the candidate safe area.
[0152] It is worth noting that, in order to improve the system's response efficiency and flexibility in the event of an emergency, the parking lot's intelligent emergency management system has introduced a dynamic mechanism of "graded safety zones." This means that there are multiple levels of general safety zones. General safety zones are temporary safety buffer zones that are temporarily designated within a local area. They are usually circular in shape and can be divided into multiple levels based on factors such as safety distance, safety redundancy, and environmental obstacles. If the estimated passage time is less than the time it takes for the vehicle to spontaneously combust, a higher-level safety zone is prioritized to isolate the target vehicle. If multiple safety zones of the same level meet the conditions, the one with the shorter estimated passage time is prioritized.
[0153] For example, when a fire risk is detected in a target vehicle, a vehicle relocation robot is invoked to go to the target vehicle, perform a multimodal vehicle status assessment, estimate the time of spontaneous combustion, and plan the best candidate route to move the target vehicle to the optimal safety zone, along with the estimated travel time for that route. If the time of spontaneous combustion of the target vehicle is longer than the estimated travel time to move the target vehicle to the optimal safety zone, the vehicle relocation robot is controlled to move the target vehicle to the optimal safety zone. Otherwise, a "leveled safety zone" mechanism is activated. Based on the current location of the vehicle relocation robot, a new route assessment and risk comparison are performed to select the emergency route with the highest safety level where the target vehicle can be successfully moved before it catches fire. The vehicle relocation robot is then controlled to move the target vehicle along this route.
[0154] For example, the delineation of graded safety zones is as follows: Centered on the target vehicle, and combined with dynamic environmental data such as the current available space, vehicle density, and facility layout of the parking lot, several grades of general safety zones are formed. Each general safety zone is assigned clear safety distance requirements and risk buffer standards. The equipment management platform refreshes the center coordinates, radius, and safety level of these zones frequently, and selects the optimal reachable general safety zone as the target area for relocation based on the hard condition that "the estimated passage time is less than the time of vehicle spontaneous combustion," thereby maximizing the success rate of handling the situation.
[0155] Optionally, the general safety zone can be divided into four levels based on the environmental safety margin and the degree of interference. The specific level definitions are as follows:
[0156] (1) Level 1 Safety Zone: This is the highest priority zone. It requires that the target vehicle maintain a distance of at least 5 meters from other vehicles after it is placed, ideally more than 10 meters. This zone should be as close as possible to the current location of the target vehicle to facilitate rapid relocation and risk isolation. At the same time, it should ensure that the path does not interfere with other static vehicles or affect key facilities in the parking lot (such as electrical boxes, cameras, anomaly detection equipment, etc.). It should also avoid major electrical equipment and fire interface points.
[0157] (2) Level 2 Safety Zone: Applicable to scenarios with slightly limited space. The minimum distance between the target vehicle and surrounding vehicles should not be less than 3 meters, and it is recommended to maintain a distance of more than 5 meters. It should also avoid major electrical equipment and fire interface points. This zone improves the deployable coverage while maintaining a high level of security.
[0158] (3) Level 3 Safety Zone: As an optional alternative, this is suitable for more compact areas, requiring a minimum distance of 1 meter from surrounding vehicles, with a recommended distance of 3 meters or more. This type of zone has basic thermal radiation isolation capabilities and can be quickly deployed in most underground parking garage passageways or multi-parking combination spaces.
[0159] (4) Level 4 Safety Zone: This is the lowest level of emergency parking area, suitable for situations where space is extremely limited and other areas are unavailable. This zone requires the target vehicle to maintain a distance of at least 0.5 meters from surrounding vehicles, and preferably more than 1 meter. It is suitable for extreme scenarios where every second counts and the margin for error is minimal.
[0160] It is worth noting that the classification of general safety zones is not limited to the specific examples mentioned above, and can be divided according to actual needs, which is not limited here. During vehicle movement, the path can be dynamically adjusted in real time, such as when a fire source is detected to spread or the path is blocked. During vehicle movement, two termination conditions need to be continuously monitored: (1) receiving a cancellation instruction from the superior platform (such as a cancellation instruction generated due to a false alarm or changes in on-site conditions); (2) the vehicle moving robot completes the predetermined vehicle moving task. After either condition is met, the vehicle moving robot will safely and smoothly place the target vehicle in the selected safety zone and automatically return to its original position or standby area in order to perform subsequent tasks or wait for the intervention of fire rescue forces.
[0161] Optionally, the target vehicle can start automatically and move to the corresponding safe area according to the target path, or the vehicle can be moved to the corresponding safe area by controlling the vehicle-moving robot according to the target path. The following is a brief introduction to the working principle of the vehicle-moving robot: One of the core functional components of the vehicle-moving robot is a tire clamping mechanism with automatic positioning, clamping, and lifting functions. This mechanism can accurately align, clamp, and lift the vehicle's tires after the robot body moves under the vehicle to be moved, thereby achieving smooth lifting and preparation for movement. The vehicle-moving robot includes a movable chassis with retractable guide arms and a sensing system at its front and rear ends. After receiving the vehicle's position information, the robot activates the positioning module, controlling itself to move along the ground to the center area under the vehicle, and automatically identifies and models the positions of the four tires through a combination of visual sensors, LiDAR, etc. After completing tire positioning, the robot activates the tire clamping mechanisms located at the four corners. Each tire clamping mechanism includes: (1) a telescopic arm: driven by an electric push rod or hydraulic cylinder, which can extend horizontally from inside the robot host to both sides of the tire; (2) a gripper structure: the grippers are made of high-strength metal material, and the inner wall is covered with rubber pads to avoid damaging the tire. The shape of the grippers matches the curvature of the outer edge of the tire, which can achieve stable coverage; (3) a synchronous control system: the four sets of grippers execute clamping actions synchronously through the central control system to ensure uniform force during clamping and prevent the vehicle from tilting; (4) a lifting mechanism: each set of grippers is equipped with an independent lifting device (such as a screw jack or scissor lift platform), which slowly lifts the vehicle after clamping and raises the vehicle to a suspended state. The height from the ground is generally set to 5-10cm to facilitate subsequent movement.
[0162] It is worth noting that the structure of the car-moving robot is not limited to the specific structure mentioned above, and can be set according to the actual situation, which is not limited here.
[0163] Optionally, the method for determining the risk of spontaneous combustion includes at least one of the following: 1. Using sensor devices installed in the parking lot to detect the real-time monitoring data of the target vehicle and judging whether the target vehicle has a risk of spontaneous combustion based on the data (such as a visible light camera detecting flames, an infrared thermal imaging camera detecting that the temperature of the target vehicle exceeds the set temperature threshold, or a smoke detector detecting combustion smoke, etc.). 2. Using the Battery Management System (BMS) equipped in current mainstream vehicles to monitor the key parameters (temperature, voltage, current, etc.) of the vehicle's power battery, and using the onboard intelligent computing unit to perform multi-source data fusion analysis to achieve early identification of potential safety risks such as thermal runaway (i.e., the risk of spontaneous combustion). Once it is determined that the vehicle is in a pre-combustion risk state, the vehicle can immediately trigger the sound and light alarm device (including buzzer, warning light, dashboard warning icon, etc.) to remind the surrounding personnel to respond in time. In order to achieve efficient information linkage, once the vehicle is determined to have a risk of spontaneous combustion, the warning information can be automatically sent to: (1) the local gateway corresponding to the equipment management platform of the parking lot intelligent fire extinguishing system; (2) the vehicle owner's mobile terminal; (3) the remote management platform. If the communication protocol supports it, the warning information will also be linked with the vehicle anomaly monitor via short-range communication links such as Bluetooth, WiFi Direct, and the parking lot's local broadcasting system, enhancing the speed of local emergency response and the multi-channel redundancy of information transmission. Given the possibility of weak or unavailable 4G / 5G signals in environments such as underground parking lots, cellular communication is designed as an auxiliary communication method to ensure communication stability and reliability. WiFi Direct is a wireless technology that allows devices to communicate directly without the need for traditional wireless routers or access points. 3. Through Figure 2 The data fusion method shown determines whether a vehicle has a risk of spontaneous combustion. Then, when the vehicle-moving robot moves to the location of a vehicle with a risk of spontaneous combustion, it also uses this method. Figure 2 The data fusion method shown is used to assess the current state of the vehicle.
[0164] Preferably, after determining that the target vehicle has a risk of spontaneous combustion, the vehicle relocation robot is controlled to move to the target vehicle, and the sensor devices installed on the vehicle relocation robot are used to conduct a relatively comprehensive monitoring of the target vehicle at close range. The real-time monitoring data obtained by the sensor devices on the vehicle relocation robot is used to estimate the vehicle condition and predict the time of spontaneous combustion. If necessary, the vehicle condition can be estimated and the time of spontaneous combustion can be predicted by combining the sensor devices of the parking lot, the target vehicle and the vehicle relocation robot.
[0165] It is worth noting that with the rapid growth in the number of new energy vehicles, the safety hazards posed by the thermal runaway of their power batteries are becoming a major challenge for urban parking management, especially indoor and underground parking lots. Once lithium-ion batteries experience thermal runaway, they are often accompanied by rapid combustion, reignition, or even explosion. This is drastically different from the emergency response to fires in traditional fuel vehicles, and existing fire prevention measures are insufficient to effectively address such complex and sudden safety events. Underground parking lots, in particular, are characterized by enclosed spaces, poor ventilation, extremely rapid fire spread, wide smoke diffusion, and dense parking spaces with narrow passageways, making efficient manual emergency response difficult and significantly increasing the risk. The high temperatures and harmful gases generated during combustion pose a significant threat to human life and property protection, thus requiring timely isolation of vehicles at risk of spontaneous combustion. However, existing vehicle relocation or firefighting robots mostly employ fixed path planning or basic obstacle avoidance algorithms, making it difficult to assess the fire's development stage and predict the remaining combustion time in dynamic and complex environments, and to dynamically adjust task strategies accordingly. In actual operation, vehicle relocation robots often face numerous practical obstacles after receiving a fire warning, such as path obstruction, limited space, and long distances to the target location. Relying solely on a single safe zone for transport may delay the optimal handling window due to insufficient transport time, or even cause the pre-ignition vehicle to catch fire during movement, thereby damaging the robotic equipment, triggering a chain fire, greatly expanding the scope of the accident, and seriously threatening the safety of surrounding personnel and other vehicles.
[0166] Compared with existing technologies, this invention provides an intelligent fire suppression system for parking lots to better predict vehicle conditions and isolate vehicles at risk of spontaneous combustion. This system utilizes deep fusion of multimodal data to detect vehicle conditions and assists in the intelligent relocation of vehicles at risk of spontaneous combustion, aiming to improve parking lot safety management and emergency response capabilities. The system integrates multiple core modules, including a vehicle anomaly monitor, an equipment management platform, an intelligent relocation robot, a multimodal vehicle status assessment mechanism, and multi-level safety zone determination and intelligent transfer decision-making. By deploying various sensors within the parking lot, the system achieves comprehensive monitoring of vehicle status. The equipment management platform analyzes and processes real-time data, deeply fusing multimodal image data with non-image data to obtain multimodal fusion features. These features are then used to accurately assess vehicle status, improving the accuracy of vehicle condition detection and enabling timely detection of vehicle anomalies. If an anomaly is detected, the equipment management platform quickly determines the location of the abnormal vehicle, provides this information to the relocation robot, and notifies on-duty personnel via an alarm system. In the subsequent handling process, the vehicle relocation robot, based on dynamic path planning and multimodal data evaluation, intelligently assesses the fire risk and safe relocation needs of the vehicles. Through a dual judgment mechanism of "optimal safety zone + multi-level safety zone" and intelligent scheduling strategy, the system can select the optimal transfer route and parking location to minimize potential dangers. Furthermore, by constructing a closed loop of "perception-decision-execution-feedback," the system effectively improves the intelligence level and reliability of emergency response to fires involving new energy vehicles. Simultaneously, through the effective management of multi-level safety zones, it enhances the overall safety of the parking lot, demonstrating the application potential of multimodal data fusion in intelligent transportation and safety management.
[0167] See Figure 4 , Figure 4 This is a schematic diagram of a vehicle condition prediction device based on multimodal data fusion provided in an embodiment of the present invention. The vehicle condition prediction device 20 based on multimodal data fusion includes:
[0168] The data acquisition module 21 is used to acquire real-time monitoring data of the vehicle; wherein, the real-time monitoring data includes non-image data and image data in at least two modalities;
[0169] Feature extraction module 22 is used to perform feature extraction and encoding as well as dimension unification processing operations on the non-image data and the image data respectively to obtain non-image features and image features;
[0170] The location embedding module 23 is used to embed spatial location information into the image features to generate location-aware features;
[0171] The joint modeling module 24 is used to jointly model the position-aware features of the image data of each modality based on the attention mechanism and the hybrid expert module, and generate intermediate features of the image data of each modality.
[0172] Image fusion module 25 is used to fuse all the intermediate features based on a local-global network to generate image fusion features;
[0173] Feature fusion module 26 is used to fuse the image fusion features and the non-image features to generate multimodal fusion features;
[0174] The vehicle condition prediction module 27 is used to predict the vehicle condition prediction result based on multimodal fusion features.
[0175] It is worth noting that the specific working process of the vehicle condition prediction device based on multimodal data fusion can be referred to the working process of the vehicle condition prediction method based on multimodal data fusion described in the above embodiments, and will not be repeated here.
[0176] Compared with existing calculations, the vehicle condition prediction device based on multimodal data fusion provided in this embodiment of the invention first acquires non-image data of the vehicle and image data of at least two modalities, and performs feature extraction and encoding, as well as dimensionality unification processing, on the non-image data and the image data respectively to obtain non-image features and image features; then, spatial location information is embedded into the image features to generate location-aware features, and the location-aware features of the image data of each modality are jointly modeled based on an attention mechanism and a hybrid expert module to generate intermediate features of the image data of each modality; next, all the intermediate features are fused based on a local-global network to generate image fusion features; the image fusion features and the non-image features are fused to generate multimodal fusion features; finally, the vehicle condition prediction result is predicted based on the multimodal fusion features. Therefore, this embodiment of the invention can obtain multimodal fusion features by deeply fusing image data of multiple modalities and then fusing it with non-image data, and use the multimodal fusion features to accurately assess the vehicle condition, improving the accuracy of vehicle condition detection and enabling timely detection of vehicle anomalies.
[0177] See Figure 5 , Figure 5 This is a schematic diagram of a vehicle condition prediction device based on multimodal data fusion provided in an embodiment of the present invention. The vehicle condition prediction device 30 based on multimodal data fusion includes a processor 31, a memory 32, and a computer program stored in the memory 32 and configured to be executed by the processor 31. When the processor 31 executes the computer program, it implements the steps as described in the above embodiment of the vehicle condition prediction method based on multimodal data fusion, for example... Figure 1The steps S1 to S7 described above; or, when the processor 31 executes the computer program, it implements the functions of each module in the above-described device embodiments.
[0178] For example, the computer program can be divided into one or more modules, which are stored in the memory 32 and executed by the processor 31 to complete the present invention. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the vehicle condition prediction device based on multimodal data fusion. For example, the computer program can be divided into multiple modules. The specific working process of each module can be referred to the working process of the vehicle condition prediction device based on multimodal data fusion described in the above embodiments, and will not be repeated here.
[0179] The vehicle condition prediction device based on multimodal data fusion can be a computing device such as a desktop computer, laptop, handheld computer, or cloud server. The device may include, but is not limited to, a processor 31 and a memory 32. Those skilled in the art will understand that the device may also include input / output devices, network access devices, buses, etc.
[0180] The processor 31 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. The processor 31 is the control center of the vehicle condition prediction device based on multimodal data fusion, connecting all parts of the device via various interfaces and lines.
[0181] The memory 32 can be used to store the computer program and / or modules. The processor 31 implements various functions of the vehicle condition prediction device based on multimodal data fusion by running or executing the computer program and / or modules stored in the memory 32 and calling the data stored in the memory 32. The memory 32 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function, etc.; the data storage area may store data created based on the use of the mobile phone, etc. In addition, the memory 32 may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0182] The module integrated into the vehicle condition prediction device based on multimodal data fusion, if implemented as a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by the processor 31, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0183] This invention also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the vehicle condition prediction method based on multimodal data fusion as described in any of the above embodiments.
[0184] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A vehicle condition prediction method based on multimodal data fusion, characterized in that, include: Acquire real-time monitoring data of the vehicle; wherein the real-time monitoring data includes non-image data and image data in at least two modalities; Feature extraction and encoding, as well as dimension unification processing, are performed on the non-image data and the image data respectively to obtain non-image features and image features; Spatial location information is embedded into the image features to generate location-aware features; The position-aware features of the image data of each modality are jointly modeled based on the attention mechanism and the hybrid expert module to generate intermediate features of the image data of each modality; Based on a local-global network, all the intermediate features are fused to generate image fusion features; The image fusion features and the non-image features are fused to generate multimodal fusion features; The vehicle condition prediction result is predicted based on multimodal fusion features; The attention mechanism and hybrid expert module jointly model the position-aware features of the image data of each modality to generate intermediate features of the image data of each modality, including: For the image data of the first modality, a first query vector is generated based on the first location-aware feature, and a first key vector and a first value vector are generated based on the second location-aware feature; wherein, the first location-aware feature is the location-aware data of the image data of the first modality, and the image data of the first modality is one of the image data of all modalities; Based on the attention mechanism, a first multi-attention output is generated according to the first query vector, the first key vector, and the first value vector; After performing a residual connection between the first multi-attention output and the first position-aware feature, a layer normalization operation is performed to obtain the first layer normalized output. The first layer normalized output is input into the hybrid expert module for processing to obtain the first hybrid expert output. After performing residual processing on the first hybrid expert output and the first layer normalized output, a layer normalization operation is performed to obtain the first intermediate feature; The non-image data is on-site smoke information, and the image data includes infrared images and visible light images.
2. The vehicle condition prediction method based on multimodal data fusion as described in claim 1, characterized in that, The process of fusing all intermediate features based on a local-global network to generate image fusion features includes: By concatenating all the intermediate features, a combined feature is obtained; The combined features are processed using a local-global network to extract multi-scale features; Image fusion features are generated based on the multi-scale features and all the intermediate features.
3. The vehicle condition prediction method based on multimodal data fusion as described in claim 1, characterized in that, The step of fusing the image fusion features and the non-image features to generate multimodal fusion features includes: The image fusion features and the non-image features are concatenated to generate concatenated features; The contextual temporal information of the spliced features is extracted, and image modality weights and non-image modality weights are learned based on the contextual temporal information. The image fusion features and the non-image features are then weighted and summed based on the image modality weights and the non-image modality weights to generate weighted fusion features. The splicing features are subjected to a nonlinear transformation to generate a residual vector; Multimodal fusion features are generated based on the contextual temporal information, the weighted fusion features, and the residual vector.
4. The vehicle condition prediction method based on multimodal data fusion as described in any one of claims 1 to 3, characterized in that, Also includes: If the vehicle condition assessment result includes the time of vehicle spontaneous combustion, then the vehicle is the target vehicle; Obtain the current vehicle location of the target vehicle in the parking lot, the location of the safe zone of the parking lot, the road layout information of the parking lot, and the information of key equipment on both sides of the road; Starting from the current vehicle location and ending at the location of the safe zone, path planning is performed based on the road layout information to obtain at least one candidate path and its estimated travel time. The estimated travel time is limited to be less than the time of vehicle spontaneous combustion, and a target route is selected from the candidate routes; The target vehicle is moved based on the target path.
5. A vehicle condition prediction device based on multimodal data fusion, characterized in that, include: The data acquisition module is used to acquire real-time monitoring data of the vehicle; wherein, the real-time monitoring data includes non-image data and image data in at least two modalities; The feature extraction module is used to perform feature extraction and encoding as well as dimension unification processing on the non-image data and the image data respectively, to obtain non-image features and image features; The location embedding module is used to embed spatial location information into the image features to generate location-aware features; The joint modeling module is used to jointly model the position-aware features of the image data of each modality based on the attention mechanism and the hybrid expert module, and generate intermediate features of the image data of each modality. An image fusion module is used to fuse all the intermediate features based on a local-global network to generate image fusion features; The feature fusion module is used to fuse the image fusion features and the non-image features to generate multimodal fusion features; The vehicle condition prediction module is used to predict the vehicle condition prediction result based on multimodal fusion features; The joint modeling module is specifically used for: For the image data of the first modality, a first query vector is generated based on the first location-aware feature, and a first key vector and a first value vector are generated based on the second location-aware feature; wherein, the first location-aware feature is the location-aware data of the image data of the first modality, and the image data of the first modality is one of the image data of all modalities; Based on the attention mechanism, a first multi-attention output is generated according to the first query vector, the first key vector, and the first value vector; After performing a residual connection between the first multi-attention output and the first position-aware feature, a layer normalization operation is performed to obtain the first layer normalized output. The first layer normalized output is input into the hybrid expert module for processing to obtain the first hybrid expert output. After performing residual processing on the first hybrid expert output and the first layer normalized output, a layer normalization operation is performed to obtain the first intermediate feature; The non-image data is on-site smoke information, and the image data includes infrared images and visible light images.
6. A vehicle condition prediction device based on multimodal data fusion, characterized in that, The system includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the vehicle condition prediction method based on multimodal data fusion as described in any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device on which the computer-readable storage medium is located to perform the vehicle condition prediction method based on multimodal data fusion as described in any one of claims 1 to 4.
8. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the vehicle condition prediction method based on multimodal data fusion as described in any one of claims 1 to 4.