Bridge deck ice and snow early warning method and system based on neural network and multi-modal data fusion
Through the bridge deck ice and snow warning method that integrates neural networks and multimodal data, combined with images and meteorological parameters, the ice thickness is predicted and the optimal de-icing plan is selected, which solves the problems of insufficient recognition accuracy and decision-making in the bridge deck ice and snow warning system, and achieves efficient and environmentally friendly de-icing effects.
Patent Information
- Application Number
- CN202510616065.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-05-14
Smart Images

Figure CN120597191A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of highway engineering technology, and specifically relates to a bridge deck ice and snow early warning method and system based on the fusion of neural network and multimodal data. Background Art
[0002] Among common accidents in the transportation sector, traffic accidents and economic losses occur more frequently in adverse weather conditions (such as rain, snow, and severe cold), seriously impacting vehicle safety and road operation efficiency. In snowy and icy weather, with low temperatures and high humidity, moisture or ice mist in the air easily condenses on the road surface, forming a layer of ice. Bridge decks, due to faster ventilation and heat dissipation, typically drop below freezing more quickly than ordinary roads, resulting in faster ice formation and longer ice retention. In some environments, when temperatures drop again at night, melted or partially melted ice and snow can easily refreeze, further increasing the slipperiness of the road.
[0003] Studies have shown that icy roads significantly reduce the coefficient of friction between tires and the road: from approximately 0.6 on dry asphalt, to approximately 0.4 on rainy days, to approximately 0.28 on snowy roads, and even as low as 0.18 on icy roads, significantly increasing vehicle braking distances. When ice forms on a bridge surface, vehicle adhesion and maneuverability are significantly reduced. Excessive speed or insufficient distance between vehicles can lead to accidents such as skidding and rear-end collisions. While some vehicles improve grip by installing snow chains or using winter tires, it is still difficult to completely eliminate the risk of accidents on severely icy bridge surfaces.
[0004] Existing solutions for dealing with snow and ice on bridges in winter typically involve spreading salt, spraying chemical deicing agents, or mechanical snow removal. However, timely and accurate assessment of bridge surface ice and snow conditions to determine when to initiate deicing measures and the appropriate deicing solution for different environmental conditions remain significant challenges. In recent years, some machine vision or meteorological sensing methods have been gradually applied to ice and snow warning systems, but actual deployment still faces the following major shortcomings:
[0005] Traditional solutions usually only rely on temperature, humidity or simple meteorological parameters, ignoring the integration of multimodal information such as visual texture, wind speed changes, and road friction coefficient. This results in insufficient accuracy and timeliness in the recognition of ice and snow conditions, and false alarms or missed reports are relatively common.
[0006] Many systems are based on pre-set fixed thresholds or static models, which make it difficult for them to adapt to different bridge deck structures (such as steel bridges and concrete bridges) and changing environmental conditions (such as aquatic environments and mountainous environments). They are also unable to accurately model and predict seasonal and real-time meteorological changes.
[0007] When selecting a de-icing solution, people often only consider a single objective (such as safety or cost), while ignoring the comprehensive trade-offs of multiple objectives such as economy, environmental protection, and operational efficiency. This may result in waste of de-icing resources or environmental pollution, and may also reduce safety under certain conditions. Summary of the Invention
[0008] The purpose of the present invention is to address the deficiencies of the above-mentioned background technology and to provide a bridge deck ice and snow warning method and system based on the fusion of neural networks and multimodal data, which can provide accurate warning of the bridge deck icing trend, thereby reducing the accident rate and improving the efficiency of snow melting and ice removal.
[0009] The technical solution adopted by the present invention is: a bridge surface ice and snow early warning method based on neural network and multimodal data fusion, comprising the following steps:
[0010] forming a multimodal feature representation based on an image of the bridge deck to be evaluated, meteorological parameters, and a bridge deck friction coefficient; the multimodal feature representation comprising a temperature-visual feature and an ice and snow texture feature;
[0011] The multimodal feature representations are fused through a learnable dynamic weighting mechanism to obtain a unified feature embedding representation;
[0012] Inputting the time series represented by the feature embedding into a time series neural network model to predict the future ice thickness of the bridge deck;
[0013] Generate warning information based on predicted ice thickness.
[0014] The above technical solution also includes the following steps: calculating the normalized performance value of each decision objective of multiple candidate de-icing schemes based on the predicted ice thickness, and calculating the evaluation value of each candidate de-icing scheme based on the corresponding weight of the decision objective to select the optimal de-icing scheme; the multiple decision objectives include at least safety, economy and environmental protection.
[0015] The above technical solution also includes the following steps: storing the fused feature embedding representation in a vector database to construct a historical ice and snow event feature library; using an approximate nearest neighbor indexing algorithm based on a graph structure to construct a topological structure for the feature library, thereby realizing fast nearest neighbor retrieval of the current feature embedding representation; inputting the retrieved historical feature vector and the time series of the feature embedding representation into a temporal neural network model to provide a historical context reference for ice thickness prediction.
[0016] In the above technical solution, when the predicted ice thickness or bridge surface friction coefficient is lower than the corresponding set threshold, an early warning message is output.
[0017] In the above technical solution, the process of generating the temperature-visual feature includes:
[0018] The multispectral image of the bridge deck and the corresponding temperature field matrix are spliced channel by channel to form a multi-channel input, which includes at least a visible light image data channel and a temperature data channel;
[0019] The multi-channel image is first normalized and randomly masked, and then input into a pre-trained visual encoder based on a Transformer structure. The parameters of the front layer of the visual encoder are partially frozen to maintain the pre-trained visual representation, and the back layer introduces a temperature field attention mechanism to give higher weight to temperature-sensitive areas;
[0020] After processing by the visual encoder, a high-dimensional feature vector of fused temperature and visual information is output for subsequent multimodal feature fusion processing.
[0021] In the above technical solution, the process of generating the ice and snow texture features includes:
[0022] A convolutional neural network-based backbone model is used to freeze some convolutional layers or residual blocks, and to fuse bridge deck image data, meteorological data, and bridge deck friction coefficient data into inputs.
[0023] After preliminary feature extraction, a spatial pyramid pooling layer is added after the last residual block of the network. The spatial pyramid pooling layer uses multiple pooling windows of different sizes to pool the output features and concatenates the pooling results of each scale into a high-dimensional vector. After batch normalization and nonlinear activation function processing, the multidimensional ice and snow texture features are output.
[0024] In the above technical solution, the input data preprocessing step of the backbone model includes:
[0025] The bridge deck image data, meteorological data, and bridge deck friction coefficient data are stored in a unified format in a database or file system. After reading the data from the database or file system, the bridge deck image data is subjected to data enhancement processing such as random rotation and color dithering. At the same time, the read meteorological data and bridge deck friction coefficient data are matched and jointly preprocessed with the processed image data in a tabular form, so that when input into the backbone model, all modal data are unified into a standardized format.
[0026] In the above technical solution, the process of fusing the multimodal feature representations through a learnable dynamic weighting mechanism includes:
[0027] First, the cosine similarity between the temperature-visual features and the ice and snow texture features is calculated;
[0028] When the similarity reaches or exceeds the preset threshold, the original weights are directly used to perform weighted fusion on the two features to obtain the fused features;
[0029] When the similarity is lower than the preset threshold, the weight of the temperature-visual feature is reduced according to the difference between the cosine similarity and the threshold, and the weight of the ice and snow texture feature is increased accordingly.
[0030] Finally, the adjusted weights are used to perform weighted fusion on the two features to obtain the final fusion features.
[0031] In the above technical solution, the fuzzy hierarchical analysis method decision model constructs a fuzzy judgment matrix and uses the fuzzy characteristic root method to calculate the comprehensive weight of each decision target; based on the predicted ice thickness, the weights of multiple decision targets in the fuzzy hierarchical analysis method decision model are dynamically adjusted.
[0032] The present invention also provides a bridge surface ice and snow warning system based on neural network and multimodal data fusion, which is used to implement the method described in the above technical solution, including:
[0033] A multimodal feature generation module, configured to generate a multimodal feature representation based on an image of the bridge deck to be evaluated, meteorological parameters, and a bridge deck friction coefficient; the multimodal feature representation includes temperature-visual features and ice and snow texture features;
[0034] A multimodal feature fusion module, configured to fuse the multimodal feature representations through a learnable dynamic weighting mechanism to obtain a unified feature embedding representation;
[0035] An ice thickness prediction module is used to input the time series represented by the feature embedding into a time series neural network model to predict the future ice thickness of the bridge deck;
[0036] The warning generation module is used to generate warning information based on the predicted ice thickness.
[0037] A decision generation module is configured to calculate a normalized effectiveness value for each decision objective of a plurality of candidate de-icing schemes based on the predicted ice thickness, and to calculate an evaluation value for each candidate de-icing scheme based on the corresponding weights of the decision objectives, so as to select an optimal de-icing scheme; the plurality of decision objectives including at least safety, economy, and environmental protection.
[0038] The beneficial effects of the present invention are as follows: the present invention comprehensively considers image data (such as visual texture), meteorological parameters and bridge surface friction coefficient to achieve multi-angle characterization of bridge surface icing and improve prediction accuracy; it takes into account the dynamic process of bridge surface ice thickness changing over time, and is closer to the actual icing and melting mechanism in practical applications; the effective fusion of different modal information can significantly reduce problems such as false alarms and missed alarms, and provide a more reliable data basis for bridge surface ice and snow warnings; from the original multimodal data to the final output warning information, a practically deployable end-to-end solution is formed.
[0039] Furthermore, the present invention not only considers safety, but also integrates economic and environmental factors to avoid single-target decision-making that causes waste of resources or secondary pollution; it conducts quantitative evaluation of different de-icing methods to reduce blindness and decision-making errors; the ice thickness prediction results directly affect the decision-making process and can be adjusted as actual conditions change, thereby improving de-icing efficiency and scientificity.
[0040] Furthermore, the present invention compares the characteristics of historical ice and snow events with the current situation, which can quickly find similar scenarios and help the model incorporate richer prior information into the prediction; based on the approximate nearest neighbor algorithm, it can efficiently search in a large-scale feature library, and combined with the time series prediction model to improve computing efficiency; historical context reference can make up for the shortcomings of single observation data and improve the accuracy and robustness of the prediction.
[0041] Furthermore, the ice thickness and friction coefficient of the present invention can both characterize the slipperiness of the road surface, and the dual judgment can reduce missed or false alarms; the threshold judgment can trigger an immediate warning after the model output, and can distinguish the warning level according to different levels of thresholds; the prediction results are linked to actual operational safety requirements in a simple and easy-to-understand way, which is convenient for deployment and execution.
[0042] Furthermore, the present invention splices the multispectral image of the bridge deck and the temperature field matrix into a multi-channel input. The multispectral image brings richer visual information, and the temperature channel helps to highlight local thermal features and improve the ability to recognize icing on the bridge deck. It utilizes the common features of large-scale data training to effectively avoid training from scratch, shorten the training cycle and improve accuracy. In the latter layer, it focuses on temperature-sensitive areas to ensure higher detection accuracy in low-temperature and potential icing areas. Through random mask enhancement, the model has better robustness and a certain degree of fault tolerance to environmental noise or information loss.
[0043] Furthermore, the present invention retains the pre-trained network's ability to extract general visual features by freezing some convolutional layers, while reducing the difficulty of model training and the amount of parameter adjustment; it obtains features under different receptive fields through the spatial pyramid pooling layer, which can better identify the texture differences of ice and snow at different scales; the combination of batch normalization and nonlinear activation further enhances the model's learning ability, and the output ice and snow texture features have a finer resolution for iced areas.
[0044] Furthermore, the present invention uses unified storage in a database or file system, which is conducive to data consistency and batch retrieval, avoiding multi-source data disorder; image data is aligned with non-image data such as meteorology and friction coefficient, so that multimodal information can be effectively integrated before entering the network; random rotation and color jittering allow the model to adapt to different shooting angles and lighting conditions, further reducing errors in actual deployment.
[0045] Furthermore, the present invention dynamically allocates weights according to feature similarity to ensure automatic adjustment in the event of information conflict or large differences, thereby reducing fusion errors; when a certain feature is unreliable (the similarity is low), the influence of other reliable features can be moderately amplified to enhance the robustness of the model; the threshold and adjustment coefficient can be optimized during training or verification to match the fusion process with the data distribution and improve the prediction effect.
[0046] Furthermore, the present invention adopts a fuzzy judgment matrix to better characterize the ambiguity and uncertainty existing in experts or actual scenarios, and is more flexible than the traditional hierarchical analysis method; when calculating the weights of multiple objectives such as safety, economy, and environmental protection, it avoids the problem of excessive subjective bias and makes the comprehensive weight more reasonable; bridge surface icing and weather conditions are often random and ambiguous, and the introduction of the fuzzy hierarchical analysis method is closer to the actual situation and improves the scientific nature of decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a schematic diagram of the overall process of the embodiment;
[0048] Figure 2 1 is a flow chart of multimodal feature fusion of an embodiment;
[0049] Figure 3 FAHP decision matrix diagram of the embodiment. DETAILED DESCRIPTION
[0050] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments to facilitate a clear understanding of the present invention, but they do not constitute a limitation to the present invention.
[0051] Example 1
[0052] like Figure 1 As shown, the present invention provides a bridge surface ice and snow early warning method based on neural network and multimodal data fusion, comprising the following steps:
[0053] S1, forming a multimodal feature representation based on an image of the bridge deck to be evaluated, meteorological parameters, and a bridge deck friction coefficient; the multimodal feature representation includes temperature-visual features and ice and snow texture features;
[0054] S2, fusing the multimodal feature representations through a learnable dynamic weighting mechanism to obtain a unified feature embedding representation;
[0055] S3, inputting the time series represented by the feature embedding into a time series neural network model to predict the future ice thickness of the bridge deck;
[0056] S4, generating warning information based on the predicted ice thickness.
[0057] S5. Calculate the normalized effectiveness value of each decision objective of multiple candidate de-icing schemes based on the predicted ice thickness, and calculate the evaluation value of each candidate de-icing scheme based on the corresponding weight of the decision objective to select the optimal de-icing scheme; the multiple decision objectives include at least safety, economy, and environmental protection.
[0058] Specifically, in step S1, after the system is started, the data acquisition layer is synchronously activated to obtain information from various sensors on the bridge deck, including but not limited to: infrared thermal imagers, multispectral cameras, meteorological sensors (or external real-time meteorological data sources), and vibration sensors. The specific process is as follows:
[0059] When the system starts, the infrared thermal imager is first activated and its sampling frequency is set to 5 Hz, that is, 5 frames of infrared thermal images can be acquired per second.
[0060] After acquiring the original bridge deck temperature distribution from the infrared thermal imager, bicubic interpolation is used to correct the pixels, refining the thermal image's spatial resolution to approximately 0.05°C / pixel (or better). This allows for precise depiction of temperature differences across the bridge deck. The processed temperature data is then output as a frame sequence with timestamps and other identifying information, enabling subsequent matching with other sensor data.
[0061] The multispectral camera covers the visible light band (about 400-700nm) and the near-infrared band (about 850-1700nm), and synchronously collects bridge surface images at a speed of 30fps (frames per second) at 4K resolution. It can capture the reflection and absorption characteristics of snow and ice layers under different spectra.
[0062] The CLAHE algorithm is applied to the acquired multispectral images to enhance the underlying snow and ice texture. In the visible light image, the conversion to the HSV (Hue-Saturation-Value) color space helps to separate or mark the highly reflective areas of snow.
[0063] The final output is the enhanced multispectral image sequence, which can be further superimposed with the corresponding timestamp and positioning information for subsequent feature extraction.
[0064] The system can be equipped with an independent meteorological module (sensors such as temperature, humidity, wind speed, precipitation, etc.), or be connected to an external real-time meteorological data source; in this embodiment, the sensor value is read once per second by default. If the external data source has a higher frequency or comes with its own timestamp, it can also be aligned accordingly.
[0065] Continuous sensor readings such as temperature, humidity, and precipitation are smoothed using a Kalman filter to reduce occasional fluctuations or interference. For wind speed data, a sliding window of approximately 30 seconds is used to calculate variance to identify gusts (rapid, short-term increases in wind speed) to facilitate feature weighting for subsequent bridge deck icing predictions.
[0066] The output meteorological data includes wind speed (accuracy ±0.2m / s), humidity (±2%RH), precipitation (±0.1mm), etc., which can be written into the data storage module through a unified interface.
[0067] Vibration sensors distributed at key locations on the bridge deck or structure continuously collect vibration signals at a fixed sampling rate (which can be adjusted to 100 Hz or higher based on actual needs). The following formula is used to perform a Fourier transform (FFT) on the collected vibration signals to extract the power spectrum characteristics of key frequency bands such as 5-15 Hz:
[0068]
[0069] Where X(f) is the frequency domain amplitude and N is the number of sampling points. All sensor data are timestamped and stored in a ring buffer, with time synchronization accuracy within ±10ms.
[0070] This embodiment combines the dynamic and historical calibration models of bridge deck vehicle loads to obtain the equivalent friction coefficient change with an accuracy of approximately ±0.05.
[0071] The system updates the friction coefficient data to the database at set time intervals (such as every second) or event triggering (such as detecting large fluctuations), and time-aligns it with the infrared thermal image, meteorological and multispectral image at the corresponding moment.
[0072] To ensure that multimodal data accurately corresponds to the same time or the same observation window in subsequent feature fusion, the system uses a unified NTP (Network Time Protocol) or GPS timestamp for synchronization.
[0073] The data storage format of the database can be in the following forms:
[0074] Image and infrared frame data: stored in image libraries or video streams in time series, and can be indexed with metadata;
[0075] Meteorological and friction coefficient data: recorded in tabular form or in a database, and marked with accuracy or filtering status.
[0076] The collected and pre-processed multimodal data (image frames, temperature fields, meteorological parameters, friction coefficient) are uniformly stored in the database, waiting for further processing by the subsequent feature generation and fusion module.
[0077] Through the above steps, this embodiment can fully acquire multimodal information such as images of the bridge deck to be evaluated, meteorological parameters, and friction coefficients, and then perform preliminary enhancement, denoising, and time synchronization on this information, ensuring data consistency, accuracy, and real-time performance. Subsequent multimodal feature generation and fusion (such as the extraction and weighting of temperature-visual features and ice and snow texture features) can be used to perform more accurate predictive analysis based on the high-quality data output by this embodiment.
[0078] Specifically, in step S1, the process of generating the temperature-visual feature includes:
[0079] The multispectral image of the bridge deck and the corresponding temperature field matrix are spliced channel by channel to form a multi-channel input, which includes at least a visible light image data channel and a temperature data channel;
[0080] The multi-channel image is first normalized and randomly masked, and then input into a pre-trained visual encoder based on a Transformer structure. The parameters of the front layer of the visual encoder are partially frozen to maintain the pre-trained visual representation, and the back layer introduces a temperature field attention mechanism to give higher weight to temperature-sensitive areas;
[0081] After processing by the visual encoder, a high-dimensional feature vector of fused temperature and visual information is output for subsequent multimodal feature fusion processing.
[0082] Preferably, after the system is started, the multispectral image and the temperature field matrix obtained by the infrared thermal imager are channel-joined to form a multi-channel input including at least visible light image channels (R, G, B) and a temperature data channel (T). If hardware conditions permit, near-infrared or other band information can also be added to expand to more channels to further enrich the input feature dimensions. The above multi-channel image is first subjected to standardization processing such as layer normalization (LayerNorm) to ensure that the data distribution of different channels is relatively consistent, which is conducive to the subsequent network extraction of stable features.
[0083] At the same time, in order to improve the robustness and generalization ability of the model, random mask augmentation (RandomMaskAugmentation) can be performed on some areas of the input image during the training stage, that is, a certain proportion of pixels are randomly blocked or noised, thereby forcing the model to focus on the overall structure of the image and key temperature hotspots.
[0084] This embodiment uses the improved CLIP model as the main body of the visual encoder, whose structure includes a 12-layer Transformer encoding network.
[0085] Freeze the first 8 layers: To preserve the generalization capabilities and common visual features obtained through pre-training on large-scale image data, this embodiment fixes the parameters of the first 8 layers.
[0086] The temperature field attention mechanism is introduced in the last four layers: the attention weight of the temperature channel is added to the multi-head self-attention (Multi-Head Attention) in this part, that is, higher attention is allocated to areas with obvious temperature gradients, making it easier for the network to identify low-temperature potential icing areas.
[0087] The temperature field attention weight can be calculated as follows:
[0088]
[0089] Among them, T i represents the bridge deck temperature value collected by the i-th temperature sensor (unit: °C), and τ = 3.0 is the temperature scaling factor. This mechanism enables the model to give higher weight to the visual features of areas sensitive to temperature gradient changes.
[0090] After passing through the aforementioned 12-layer Transformer encoder, the model outputs a high-dimensional feature vector (e.g., 768 dimensions) that combines visible light and temperature information. Compared to traditional visual encoders that use only RGB channels, this also considers the bridge deck temperature distribution, enabling the model to recognize ice and snow with greater sensitivity and accuracy. This high-dimensional feature vector serves as the "temperature-visual feature" in the subsequent multimodal feature fusion stage, participating in the final fusion along with the "ice and snow texture feature."
[0091] Specifically, in step S1, the process of generating the ice and snow texture features includes:
[0092] A convolutional neural network-based backbone model is used to freeze some convolutional layers or residual blocks, and to fuse bridge deck image data, meteorological data, and bridge deck friction coefficient data into inputs.
[0093] After preliminary feature extraction, a spatial pyramid pooling layer is added after the last residual block of the network. The spatial pyramid pooling layer uses multiple pooling windows of different sizes to pool the output features and concatenates the pooling results of each scale into a high-dimensional vector. After batch normalization and nonlinear activation function processing, the multidimensional ice and snow texture features are output.
[0094] The input data preprocessing steps of the backbone model include:
[0095] The bridge deck image data, meteorological data, and bridge deck friction coefficient data are stored in a unified format in a database or file system. After reading the data from the database or file system, the bridge deck image data is subjected to data enhancement processing such as random rotation and color dithering. At the same time, the read meteorological data and bridge deck friction coefficient data are matched and jointly preprocessed with the processed image data in a tabular form, so that when input into the backbone model, all modal data are unified into a standardized format.
[0096] Preferably, when the system is deployed, the bridge deck image data (visible light image), meteorological data (such as temperature, humidity, wind speed, etc.) and bridge deck friction coefficient data are stored in a database or file system in a unified format.
[0097] When performing model training or online inference, the required image data is first read from the database or file system; at the same time, the meteorological data and friction coefficient data within the corresponding moment or time window are read to complete multimodal alignment in subsequent network input or fusion processing.
[0098] The bridge image data is randomly rotated within a certain angle range (e.g., ±10°) to simulate the difference in actual shooting angles and the possible perspective shift during vehicle driving. The image's hue (H), saturation (S), or brightness (V) is slightly perturbed in the HSV color space (e.g., within a range of ±0.1) to enhance the model's robustness to different lighting conditions, camera settings, and ambient light variations. Note that data augmentation is only used during training; random augmentation is disabled during inference.
[0099] The meteorological data and bridge friction coefficient data are matched in tabular form with the image data at the corresponding timestamp or frame number, achieving accurate alignment of multimodal data at the same time or within the same observation window. After alignment, a set of multimodal input samples in a standardized format (for example, image + meteorological array + friction coefficient) is formed.
[0100] After completing the above preprocessing, the image is sent as input to the ResNet50 backbone network; in order to make full use of the meteorological data and friction coefficient information, they are fused at the network input or intermediate layer through splicing or embedding, which can be determined according to the design requirements.
[0101] To preserve the basic visual features trained on large-scale general-purpose data, this implementation freezes the parameters of the first five convolutional blocks of ResNet50 (corresponding to layer indices 0-4), and only fine-tunes the subsequent layers and fully connected layers for ice and snow scenes to balance model accuracy and training cost. After passing through the last residual block of ResNet50, the SPP layer (Spatial Pyramid Pooling) is connected, and multiple pooling windows of different sizes (such as 6×6, 3×3, and 1×1) are configured to perform multi-scale pooling on the output feature map. The feature vectors obtained at different pooling scales are concatenated to obtain a high-dimensional feature description (for example, 1024 dimensions, the specific dimension is determined by the number of network channels and design requirements), which can more comprehensively depict the texture distribution of ice and snow on the bridge surface at both local and global scales.
[0102] Batch normalization (BatchNorm) is performed on the concatenated high-dimensional vectors to reduce data distribution differences between batches and improve model convergence stability. Nonlinear activation functions (such as GeLU or ReLU) are then used to further enhance the network's ability to represent complex features, ultimately outputting the final snow and ice texture features. In this example, the output is a 512-dimensional vector (this can be adjusted depending on the network design).
[0103] The ice and snow texture features obtained through the above process better reflect the local texture, edge morphology, and potential slippery area distribution of the ice and snow cover on the bridge surface; combined with the meteorological and friction coefficient information read in the step and matched with the image data, the model can more accurately capture the impact of changes in temperature, humidity, or vehicle load on ice and snow morphology when extracting local features; the features output here can be directly or through further normalization, mapping, etc., and weightedly combined with other branches (such as temperature-visual features) in the subsequent multimodal feature fusion module to provide multi-dimensional information support for ice and snow warnings or thickness predictions.
[0104] This embodiment improves the CLIP model and ResNet50 to process features with different focuses ("temperature-vision" vs. "ice and snow texture") in parallel, which can more comprehensively characterize the icing condition of the bridge surface. In both branches, the parameters of the first few layers of the pre-trained model are retained, taking into account the learning of general visual features and domain-specific features, which not only reduces training costs but also improves recognition accuracy in ice and snow scenes. The SPP layer provides a multi-scale receptive field for the ResNet50 branch, which has good detection capabilities for large-area or local features that may appear in ice and snow covered areas. Whether it is random masking, random rotation or color dithering, it can improve the robustness of the model in the face of actual complex environments and reduce the risk of overfitting. The feature vectors output by the two branches complement each other in subsequent steps (usually dynamically fused through learnable weights), helping the overall model to more accurately predict ice thickness or identify ice and snow areas.
[0105] In this embodiment, the temperature-visual feature extraction and ice and snow texture feature extraction described in step S1 are performed in parallel or alternately in the computational layer. Once completed, they are input into the feature fusion module of the next stage. This enables the formation of a multimodal, multi-scale, and multi-layered feature representation, laying a solid data foundation for the final bridge surface ice and snow warning and de-icing decision-making. Appropriate tailoring or expansion can be performed in terms of image resolution, number of channels, and CNN backbone network type to accommodate different hardware configurations or demand scenarios, while maintaining the overall concept and technical results.
[0106] Specifically, step S2 includes the following steps: first, calculating the cosine similarity between the temperature-visual feature and the ice and snow texture feature;
[0107] Dynamic weight calculation formula:
[0108] α new =α old -γ·(θ-similarity)
[0109] Where γ is the learning rate and θ is the similarity threshold.
[0110] When the similarity reaches or exceeds the preset threshold, the original weights are directly used to perform weighted fusion on the two features to obtain the fused features;
[0111] When the similarity is lower than the preset threshold, the weight of the temperature-visual feature is reduced according to the difference between the cosine similarity and the threshold, and the weight of the ice and snow texture feature is increased accordingly.
[0112] Finally, the adjusted weights are used to perform weighted fusion on the two features to obtain the final fusion features.
[0113] like Figure 2 As shown, preferably, in this embodiment, in order to fully combine the temperature-visual features (output by the improved CLIP model) and the ice and snow texture features (output by the ResNet50+SPP branch), the system introduces a dynamic feature fusion module in step S2. The specific implementation process is as follows:
[0114] Assume that the visual-temperature joint feature output by the CLIP branch is F visual ∈R 768 ; The thermodynamic (ice and snow texture) feature output by the ResNet50 branch is F thermal ∈R 512 .
[0115] To facilitate subsequent fusion, first visual and F thermal L2 normalization is performed respectively; among them, F thermal It also needs to be extended to F through a fully connected layer or other mapping layer visual The same 768-dimensional space (L2 normalization can be performed again after the fully connected layer).
[0116] Two learnable parameters λ are introduced in the fusion stage v and λ t , whose initial values are set to 0.6 and 0.4 respectively; both can be automatically optimized through back propagation, so as to gradually converge to a more suitable ratio according to different samples during the training process.
[0117] When the weight rebalancing mechanism is not triggered, the feature fusion of this embodiment can be expressed as:
[0118] F fusion =λ v ·Fvisual +λ t ·F thermal
[0119] Among them, F visual and F thermal , both are L2 normalized and mapped to a unified 768-dimensional feature space through a fully connected layer.
[0120] In order to determine the similarity or complementarity of two features, this embodiment uses cosine similarity to measure the correlation between the two features. When the similarity is lower than a threshold value θ=0.65, a weight rebalancing mechanism is triggered.
[0121] When the cosine similarity is greater than or equal to 0.65, it is considered that the temperature-visual feature and the ice and snow texture feature have a high degree of similarity or complementarity on the current sample, and the current λ is directly used. v and λ t To perform the fusion:
[0122] If the cosine similarity is less than 0.65, it means that there are certain differences or conflicts between the two features of the current sample, and the weight rebalancing mechanism needs to be triggered for dynamic adjustment.
[0123] λ' v =λ v -0.1×(0.65-S cos )
[0124] λ' t =λ t +0.1×(0.65-S cos )
[0125] Among them, S cos This is the cosine similarity calculated in real time; the preset adjustment coefficient is 0.1, which can be adjusted during training.
[0126] If the difference between the cosine similarity and its set threshold is larger, it means that the difference between the temperature-visual features and the ice and snow texture features is more obvious, and it is necessary to lean towards the ice and snow texture features to a greater extent in this training iteration.
[0127] Updated λ v and λ t The sum of the two values is approximately equal to 1, which can be fine-tuned or normalized during implementation to prevent excessive deviation of the weight value.
[0128] After the above possible rebalancing operation, the system recalculates the fused output features. This vector will be compared with the true label (such as ice thickness, ice and snow classification, etc.) in the subsequent loss function to generate a reverse gradient for updating λ v and λ t And the trainable parameters of the previous network.
[0129] Similar to traditional neural network training, this implementation uses gradient backpropagation to simultaneously update: unfrozen layer parameters in the CLIP model and ResNet50 branches; and subsequent network structures, such as the learnable weights and fully connected mapping layers in the dynamic fusion module. By iterating on large-scale training data or a limited amount of bridge surface ice and snow scene data, the multimodal feature fusion module can learn the optimal weight allocation strategy for different scenarios or environmental conditions (such as temperature differences, light intensity, and snow depth).
[0130] In this embodiment, based on cosine similarity, the model can adaptively adjust the weights of temperature-visual and ice and snow texture features to ensure that when the two features conflict, the more discriminative features can be utilized more effectively. When the similarity is high, direct fusion according to the current weight can reduce additional computational overhead and maintain stability; when the difference is large, weight rebalancing can better tap the potential value of ice and snow texture features. The learnable weights are not fixed constants, but automatically find the optimal fusion ratio suitable for the overall network performance during the training process. For ice and snow weather scenes, temperature plays a key role in the tendency to freeze, while texture features help distinguish the snow or ice cover morphology; this fusion mechanism can significantly improve the accuracy of recognition or prediction in actual bridge deck ice and snow warnings and de-icing decisions.
[0131] Through the above process, this embodiment successfully introduces cosine similarity and learnable weights during the multimodal feature fusion stage, and adaptively adjusts when the similarity falls below a threshold. Ultimately, this model possesses flexible discrimination and robustness for various bridge surface ice and snow scenarios. This feature fusion mechanism seamlessly integrates with the aforementioned temperature-visual feature generation and ice and snow texture feature generation processes, providing higher-dimensional and more reliable feature representations for subsequent time series prediction and early warning decision-making.
[0132] Specifically, step S3 also includes the following steps: storing the fused feature embedding representation in a vector database to construct a historical ice and snow event feature library; using an approximate nearest neighbor indexing algorithm based on a graph structure to construct a topological structure for the feature library, thereby realizing fast nearest neighbor retrieval of the current feature embedding representation; inputting the retrieved historical feature vector and the time series of the feature embedding representation into a temporal neural network model to provide a historical context reference for ice thickness prediction.
[0133] Preferably, multimodal features (such as temperature-visual features and ice and snow texture features) are fused through a learnable dynamic weighting mechanism to produce a unified high-dimensional feature vector. The system stores these fused feature vectors, along with the corresponding timestamp, geographic location, or other identification information (optional), in the FAISS vector database to form a historical ice and snow event feature library. This feature library is continuously updated over time or as the system collects new bridge surface ice and snow data, and its storage capacity gradually expands.
[0134] In order to quickly retrieve several records that are most similar to the current bridge deck features from the massive historical feature vectors, this embodiment uses the HNSW (Hierarchical Navigable Small World) approximate nearest neighbor indexing algorithm to construct an index topology structure.
[0135] When a new fused feature vector is inserted into the database, the FAISS engine invokes the HNSW algorithm to add the vector to the index graph and maintains adjacency based on cosine similarity or Euclidean distance between vectors (selected based on actual needs). This graph-based indexing approach significantly reduces search time while maintaining high search accuracy, ensuring system availability in real-time or near-real-time environments.
[0136] The current bridge deck ice and snow features (also derived from a feature vector obtained through multimodal fusion) are input into the FAISS vector database as a query vector. The database uses the HNSW indexing algorithm to quickly locate neighboring nodes in the high-dimensional vector space and outputs the top-K historical feature items most similar to the query vector (in this example, K is set to 5, returning the five records most similar to the current bridge deck features).
[0137] After testing, the retrieval delay of this embodiment is stable below 50ms, and the nearest neighbor retrieval can be completed within milliseconds, meeting the needs of most real-time scenarios.
[0138] The retrieved top-5 historical feature vectors, along with the time series of the current bridge deck feature vectors, are input into the time series neural network model. Leveraging the feature distributions of these historically similar scenarios, the model can better learn how ice thickness evolves under similar meteorological conditions, similar bridge deck conditions, or similar friction coefficients. This allows the model to leverage not only current data but also information from similar historical scenarios, further improving the accuracy and stability of ice thickness predictions.
[0139] Due to the efficient implementation of the HNSW approximate nearest neighbor algorithm combined with the FAISS database, massive historical features can also be retrieved in a very short time, ensuring the system response speed. Top-K similar features provide reference samples that are closer to the current bridge deck environment for time series prediction, making the neural network's learning of potential icing trends more targeted, and helping to reduce prediction errors caused by environmental differences. The FAISS vector database and HNSW index of this embodiment are not limited to bridge deck ice and snow data. If it is necessary to introduce larger-scale or different modal features (such as radar and lidar data) in the future, they can also be included in a similar way and quickly retrieved.
[0140] Specifically, in step S3, before making a prediction, the system will collect the time series of the current bridge deck features (such as the fused feature embedding representation within the last 30 minutes) and merge the retrieved Top-K historical feature vectors (and their label information) as additional input as a priori reference for "how the ice layer evolves under similar meteorological / bridge deck conditions."
[0141] There are various ways to integrate:
[0142] Directly concatenate the mean or weighted vector of historical features to the current moment features; add encoding of historical information to the initial hidden state or an intermediate layer of the time series model;
[0143] Alternatively, before the bidirectional LSTM, the retrieval history features are mapped to the same dimension as the current time series and then mixed.
[0144] This embodiment uses a bidirectional LSTM network to process time series feature sequences. The number of hidden layer units is set to 64, the input sequence length is set to 30 minutes (corresponding to 180 5Hz sampling points), and the supplementary information of the aforementioned historical feature vector is combined;
[0145] At the network output layer, the model makes multi-step predictions for future ice thickness (e.g., 15 minutes, 30 minutes), and calculates key metrics such as formation time and ice growth rate. Following the LSTM hidden layer, a fully connected layer maps the output to a scalar value representing the predicted future ice thickness (e.g., the bridge deck ice thickness 15 minutes or longer in the future).
[0146] This example uses Huber Loss instead of the traditional mean square error (MSE) or mean absolute error (MAE) to balance sensitivity to large errors with smoothness for small errors, thereby improving model convergence stability. Furthermore, regularization of the ice thickness gradient is introduced as follows:
[0147]
[0148] Where h(t) is the predicted ice thickness at time t.
[0149] During training, in addition to using Huber loss and ice thickness gradient penalty, historical feature retrieval accuracy (such as similarity distribution) can also be incorporated into the regularization target as needed to further strengthen the model's focus on high-similarity scenes.
[0150] Through back propagation, not only the bidirectional LSTM network parameters are updated, but also the historical retrieval strategy (such as thresholds, weights, etc.) can be fine-tuned, thereby continuously improving the overall prediction accuracy.
[0151] Specifically, in step S4, once the predicted future ice thickness exceeds 2 mm (set in this embodiment) or the friction coefficient is monitored to be below 0.25 in real time, an early warning is triggered;
[0152] Combined with similar historical cases, the system can provide auxiliary information such as "the risk level that may evolve to" or "how long it may take to reach the safety threshold", providing more intuitive reference for operation and maintenance personnel or vehicle drivers.
[0153] Specifically, in step S5, the fuzzy analytic hierarchy process decision model constructs a fuzzy judgment matrix and uses the fuzzy eigenvalue method to calculate the comprehensive weight of each decision target; and dynamically adjusts the weights of multiple decision targets in the fuzzy analytic hierarchy process decision model based on the predicted ice thickness.
[0154] like Figure 3 As shown, preferably, after completing the ice and snow prediction (e.g., ice thickness prediction), this embodiment uses the fuzzy analytic hierarchy process (FAHP) to construct a multi-objective decision model in order to achieve a comprehensive balance among multiple decision-making objectives such as economy, safety, and environmental protection. The main process and key implementation are as follows:
[0155] 1. Constructing three-dimensional target matrix and fuzzy judgment matrix
[0156] This system sets the three main goals of de-icing decision-making as: economy (cost, resource consumption, etc.), safety (vehicle traffic safety, accident risk reduction, etc.), and environmental protection (impact on the environment or bridge structure, emissions, etc.).
[0157] In actual scenarios, more sub-goals or detailed indicators can be expanded according to application requirements, such as snow melting efficiency, time cost, etc.
[0158] First, based on expert experience, we obtained w = (0.45, 0.37, 0.18), which correspond to the preliminary weight values of economy, safety, and environmental protection, respectively.
[0159] For each goal (such as economy vs. safety, economy vs. environmental protection, and safety vs. environmental protection), a fuzzy comparison value of relative importance is given by experts or the system experience library.
[0160] FAHP decision model constructs a three-dimensional target matrix C = [c ij ] 3×3 , where c ij Indicates the importance of the i-th goal relative to the j-th goal. The triangular fuzzy number is used to represent the elements of the judgment matrix. For example, when the expert judges that the relative importance of "economy" and "safety" is in the medium to high range, the triangular fuzzy number can be set to The triangular fuzzy numbers of safety relative to environmental protection are set to (0.4, 0.6, 0.8); the triangular fuzzy numbers of economy relative to environmental protection are set to (0.7, 0.9, 1.0).
[0161] By synthesizing, defuzzifying and normalizing the triangular fuzzy numbers, a set of relatively clear ratios is obtained for the subsequent solution of the characteristic root method.
[0162] After the fuzzy judgment matrix is converted into the corresponding fuzzy relative weight matrix, the characteristic root method (generally the maximum eigenvalue λ max and its corresponding eigenvectors) to obtain the weight vector w = (w1, w2, w3) of each target:
[0163]
[0164] To ensure that the sum of the weight vectors is 1, an appropriate normalization process can be performed to ensure that ∑w i =1.
[0165] In the previous step, the time series neural network or other prediction model has given the ice thickness prediction value y for a period of time in the future (such as 15 minutes, 30 minutes) t .
[0166] When the predicted ice thickness exceeds a certain safety threshold (such as 2 mm), the system will increase its relative attention to the "safety" goal; if the predicted ice thickness is lower, it can be more inclined to economy or environmental protection while ensuring basic safety.
[0167] Preferably, the system can set one or more thresholds and corresponding weight adjustment strategies:
[0168] w 安全性 ←w 安全性 +Δ w ×(y t -threshold)
[0169] where Δw is an adjustable coefficient; threshold represents the safety critical value. Similarly, the weight of economic or environmental goals can be reduced accordingly to maintain overall harmony.
[0170] In this way, when the predicted ice thickness is detected to be at a dangerous level, the weight of "safety" will be moderately amplified; conversely, "economy" or "environmental protection" may occupy a larger proportion.
[0171] Normalize the results of the above dynamic weight adjustment process again to get the updated final weight vector w i =(w1′,w2′,w3′), which is used to generate the next multi-objective decision evaluation. If other real-time information (such as gusts of wind or sudden temperature drops) needs to be considered, the weights can be fine-tuned simultaneously during this process to achieve more refined dynamic decision-making.
[0172] For the various deicing or snow melting schemes pre-defined by the system (such as mechanical deicing, deicing with snow melting agents, microwave deicing, etc.), the normalized performance values under the three goals of economy, safety, and environmental protection are calculated respectively.
[0173] In this embodiment, the system maintains a solution library that contains several commonly used or feasible de-icing / snow melting technical means, such as:
[0174] Mechanical de-icing: The use of mechanical equipment to remove snow and ice;
[0175] Deicing with de-icing agents (salt de-icing): spraying or spreading chemicals (salt, calcium chloride, etc.) to accelerate the melting of ice and snow;
[0176] Microwave deicing (hot water spraying): Melts or softens the ice layer through electromagnetic wave heating, with mechanical assistance.
[0177] Based on the three major goals of the system to achieve, namely, economy, safety, and environmental protection, a set of evaluation indicators is set to quantify the performance of each solution in terms of the goal. For example:
[0178] Economic indicators include:
[0179] Direct cost (Cost): de-icing cost per unit time or unit area, including equipment rental fees, chemicals, labor costs, etc.
[0180] Cost=BaseCost+k ice ×y t ,
[0181] Among them, y t is the predicted ice thickness (e.g. mm), BaseCos is the startup or fixed cost, k iceis a variable cost coefficient corresponding to ice thickness (which can be derived from historical operation records or equipment manuals).
[0182] Time cost (Time): The total time (minutes or hours) from the start to the completion of the de-icing operation, which can be converted into indirect economic losses or operational impacts.
[0183] Time=Tsetup+r×y t ,
[0184] Where TsetupT is the preparation time (machine installation, personnel in place, etc.), and r is the thickness of the ice layer that can be melted by the equipment or operating unit within 1 hour (or minute) (mm / h or mm / min).
[0185] Safety indicators include:
[0186] Residual Ice: The maximum ice thickness (mm) or amount of snow that may remain on the bridge deck after de-icing operations are completed;
[0187] If we assume that in the case of high thickness, a single operation may not be able to completely remove all ice, we can estimate it according to the empirical model:
[0188] ResidualIce=max(0,y t -E(Method,t)),,
[0189] Where E(Method, t) represents the thickness of the ice layer that can be removed after time t (or one operation cycle) under the specified deicing method Method. t If the ice capacity is large and the de-icing efficiency is limited, ResidualIce may be significantly greater than 0.
[0190] Traffic accident risk coefficient (Risk): Based on historical accident statistics or expert scores, it comprehensively measures the safety and reliability of different de-icing schemes during operation (larger values represent higher risks).
[0191] Generally speaking, the thicker the ice, the more slippery the road, and the higher the risk of vehicle driving. A related function can be used to map the "predicted ice thickness" to the "accident risk coefficient" range, such as:
[0192] Risk(y t )=y t / (y t +c),
[0193] Where c is a constant (such as 1mm or other reference value), when y t When >>c, Risk approaches 1 (high risk); otherwise, it approaches 0 (low risk).
[0194] Environmental indicators:
[0195] Chemical Emission: For each type of de-icing agent, estimate its potential pollution or corrosion to the environment (rivers, land);
[0196] When using a deicing agent type de-icing solution, the amount of de-icing agent delivered is generally proportional to the total thickness of the ice layer that needs to be melted and the area covered. Example formula:
[0197] ChemicalEmi=α*A*y t ,
[0198] Where A is the bridge deck operating area, and α represents the chemical dosage and discharge coefficient required per unit ice thickness. The value of α will vary depending on the type of chemical used or the method of delivery. Higher predicted ice thickness generally requires more chemical.
[0199] Energy consumption: such as the power demand for microwave de-icing or the fuel consumption of mechanical equipment;
[0200] For example, microwave de-icing or mechanical equipment (fuel, electricity) will run longer and use higher power when the thickness is greater, and energy consumption will increase accordingly. This can be described using linear or nonlinear functions:
[0201] Energy=P×T(y t )
[0202] Where P is the equipment power or power factor, T(y t ) is the operation duration calculated based on ice thickness.
[0203] Byproduct Impact: For example, the degree of corrosion of the metal structure of the bridge or the damage to the surrounding plants caused by the residue after spraying snow-melting agents.
[0204] As the ice layer thickens, more chemical reagents or more frequent operations may be required, resulting in more by-products such as "waste liquid" or "metal corrosion products." If historical data or experiments show that each millimeter of ice corresponds to a certain amount of corrosion or contamination residue, the proportion can be estimated relatively directly based on the ice thickness:
[0205] ByproductImpact=β*y t
[0206] Where β is the amount of by-products or environmental damage coefficient corresponding to "unit thickness of ice layer".
[0207] To facilitate comprehensive comparison, the above original indicators need to be dimensionalized and normalized to obtain a performance value (or score) between [0,1]. The larger the value, the better the performance under the goal.
[0208] The normalized formula for economic objectives is as follows:
[0209]
[0210] Among them, Cost (k) is the unit deicing cost of scheme k, Time (k) is the time required (after appropriate normalization or logarithmization), w C and w T is the weight of each sub-indicator under this goal.
[0211] The normalized formula for security objectives is as follows:
[0212]
[0213] Among them, Norm(·) represents the normalization function that maps the original value to [0,1], w res and w risk ResidualIce is the amount of residual ice and the risk factor. (k) Indicates the amount of residual ice, Risk (k) Represents the risk factor.
[0214] The normalized formula for environmental protection objectives is as follows:
[0215]
[0216] Among them, Chemical (k) Indicates the chemical pollution emission intensity of scheme k, Energy (k) Is the energy consumption per unit time or unit area, ByproductScore (k) The impact of by-products can be scored [0,1] by experts or historical data, w chem ,w energy ,w byproduct For each weight.
[0217] The final comprehensive score S for each solution k k It can be expressed as:
[0218]
[0219] in, is the normalized performance value of the k-th solution under the i-th objective.
[0220] After calculating the comprehensive scores for all candidate snow melting solutions, they are ranked from highest to lowest. Based on years of practical experience and accumulated data, this example screens out the following three de-icing solutions and scores them in terms of economy, safety, and environmental protection (the scores range from 0 to 1, with higher scores indicating a better solution in that indicator):
[0221] Mechanical deicing: Economy: 0.82; Safety: 0.75; Environmental protection: 0.63
[0222] Salt snow melting: Economic efficiency: 0.91; Safety: 0.68; Environmental protection: 0.42
[0223] Hot water spraying: Economy: 0.73; Safety: 0.85; Environmental protection: 0.79
[0224] Based on the weighted average method, a comprehensive score is calculated for each plan. Substitute the weights and the indicator scores of each plan into the calculation:
[0225] Mechanical de-icing:
[0226] 0.45×0.82+0.37×0.75+0.18×0.63≈0.78
[0227] Salt snowmelt:
[0228] 0.45×0.91+0.37×0.68+0.18×0.42≈0.72
[0229] Hot water spray:
[0230] 0.45×0.73+0.37×0.85+0.18×0.79≈0.81
[0231] Combined with the scoring results, the hot water spraying solution achieved a relatively better balance between economy, safety and environmental protection, and had the highest comprehensive score. Therefore, in this embodiment, "hot water spraying" should be used as the preferred method for bridge de-icing.
[0232] If the difference in the comprehensive score is within a certain range (such as less than 0.05), the system can select a multi-scheme collaborative operation mode and combine multiple means to accelerate de-icing; otherwise, the system directly dispatches the single scheme with the highest score or a few optimal schemes for rapid execution.
[0233] Once the decision result is generated, the corresponding actuator (such as mechanical de-icing equipment, automatic snow-melting agent spreading device, microwave preheating device, etc.) can be triggered through the system's decision engine; during the de-icing operation, the friction coefficient and ice thickness changes are monitored in real time, and the new observation values are fed back to the system to continue to correct or deactivate the decision process, forming a closed-loop control.
[0234] The system's FAHP decision-making model incorporates dynamic weight adjustment and can be applied to different bridge types, including steel and concrete bridges, as well as diverse environmental conditions, including mountainous areas, water bodies, and urban areas. Statistics show that the system's ice prediction error is controllable to ±1.2mm (95% confidence level) within the temperature range of -15°C to 5°C. Combined with FAHP dynamic decision-making, it can generate a snowmelt plan within 8 seconds, significantly more efficient than traditional methods.
[0235] In step S5, this embodiment closely integrates the fuzzy analytic hierarchy process (FAHP) decision model with the real-time predicted ice thickness. A fuzzy judgment matrix and fuzzy eigenvalue method are used to calculate the basic weight vector, which is then dynamically adjusted when the risk of bridge deck icing is detected. Ultimately, the system comprehensively scores various snowmelt solutions based on the updated multi-objective weights and selects the one with the highest priority or implements a collaborative operation strategy. This balances multiple indicators, including safety, economy, and environmental protection, significantly improving the intelligence and effectiveness of bridge deck ice and snow warning and de-icing decisions.
[0236] Example 2
[0237] The present invention also provides a bridge surface ice and snow warning system based on neural network and multimodal data fusion, which is used to implement the method described in the above technical solution, including:
[0238] A multimodal feature generation module, configured to generate a multimodal feature representation based on an image of the bridge deck to be evaluated, meteorological parameters, and a bridge deck friction coefficient; the multimodal feature representation includes temperature-visual features and ice and snow texture features;
[0239] A multimodal feature fusion module, configured to fuse the multimodal feature representations through a learnable dynamic weighting mechanism to obtain a unified feature embedding representation;
[0240] An ice thickness prediction module is used to input the time series represented by the feature embedding into a time series neural network model to predict the future ice thickness of the bridge deck;
[0241] The warning generation module is used to generate warning information based on the predicted ice thickness.
[0242] A decision generation module is configured to calculate a normalized effectiveness value for each decision objective of a plurality of candidate de-icing schemes based on the predicted ice thickness, and to calculate an evaluation value for each candidate de-icing scheme based on the corresponding weights of the decision objectives, so as to select an optimal de-icing scheme; the plurality of decision objectives including at least safety, economy, and environmental protection.
[0243] To test the system's performance during its initial deployment (without extensive historical data), this example selected three representative bridges in northern China (two steel bridges and one concrete bridge). The system monitored these bridges for 20 consecutive days in winter and collected the following multimodal data:
[0244] Visual data: A high dynamic range (HDR) camera (resolution 4096×2160, frame rate 30fps) is deployed to simultaneously acquire visible light (400-700nm) and shortwave infrared (900-1700nm) images. Approximately 800 images are acquired daily and stored on edge servers for subsequent processing.
[0245] Temperature data: Distributed fiber optic temperature sensors with an accuracy of ±0.05°C, a spatial resolution of 0.1m, and a sampling frequency of 10Hz are used to monitor the temperature gradient of the bridge deck in real time. Sensors are evenly distributed along the bridge deck to detect temperature changes in different areas.
[0246] Meteorological data: Wind speed (accuracy ±0.1m / s), relative humidity (±1%RH), and precipitation intensity (±0.05mm / h) are collected from local weather stations and transmitted to edge computing nodes using LoRa wireless technology.
[0247] It is integrated with temperature data and visual data after timestamp alignment to form a multimodal input source.
[0248] The continuous monitoring period of this embodiment is 20 days, and data is collected at different times every day (including nighttime and extreme weather conditions) to provide diverse scenarios for subsequent model training and testing.
[0249] Two certified engineers used the majority voting method to annotate the collected images and corresponding sensor data, focusing on the type and thickness of ice and snow;
[0250] Three types of ice and snow forms are defined:
[0251] Frost ice: thickness ≤1mm, granular surface;
[0252] Compacted snow: thickness 3-5mm, density >0.3g / cm 3 ;
[0253] Mixed state: The ice layer is mixed with salt particles or anti-slip materials and is distributed non-homogeneously.
[0254] The following data augmentation methods are used:
[0255] Cross-modal enhancement: A random temperature offset (±1.5°C) is applied to the infrared image to simulate the potential measurement deviation of the sensor under extreme conditions.
[0256] Spatiotemporal alignment: The optical flow algorithm is used to compensate for the spatial misalignment of the camera and temperature sensor in their installation positions, and time synchronization is performed (accuracy ≤ 5ms) to ensure that different modal data correspond to the same target area and time.
[0257] The basic training set of this embodiment includes 300 data sets (120 sets of frost ice, 100 sets of compacted snow, and 80 sets of mixed states); the extreme test set includes 50 data sets (10 sets of inversion layer ice, and 20 sets of ice and snow morphology variations after traffic rolling).
[0258] Through the above steps, a small-scale but scenario-rich set of labeled data is obtained in the initial deployment phase, providing a basis for testing the cold start performance of the system.
[0259] In an edge computing environment, an extreme test set (50 sets of data in total) was injected into the system frame by frame to simulate the actual operation process. Two typical scenarios were designed for key verification:
[0260] Scenario A: Ice forms in an inversion layer on the surface of a steel bridge: The ambient temperature gradually rises from -5°C to 0°C, and an inversion layer forms locally on the bridge deck. The system's detection accuracy and response speed are observed under rapidly changing temperature gradients.
[0261] Scenario B, variation in the compacted snow morphology after rush hour: Due to the high-frequency rolling of vehicles, the surface friction coefficient suddenly changes, and the snow layer gradually becomes compacted and mixed with mud and sand; the focus is on examining the system's ability to recognize complex snow surface textures and friction coefficient anomalies.
[0262] This example uses the following evaluation indicators:
[0263] First-frame recognition accuracy: The system's recognition accuracy when it first detects signs of ice and snow (correct classification rate for frost ice, compacted snow, and mixed states);
[0264] Steady-state error: the root mean square error of ice thickness prediction during the continuous monitoring phase (unit: mm);
[0265] False alarm decay rate: The ratio of the number of false alarms to decrease due to model updates or threshold adjustments after 24 hours of self-correction.
[0266] Cross-material consistency: The difference in test results for similar ice and snow conditions for steel and concrete bridges (the smaller the difference, the better the consistency);
[0267] Anti-interference ability: The suppression rate of interference factors such as headlight reflections and bridge cracks, that is, the difference in recognition accuracy with and without interference conditions.
[0268] Then perform cold start performance optimization:
[0269] The subsystem sets up a double buffer storage area, including: storing high-confidence positive samples (typical ice and snow features) and storing difficult negative samples (easily confused features, such as cracks, oil stains, etc.).
[0270] A dynamic threshold is introduced in the update rule. When the feature similarity is higher or lower than a certain standard, the sample is automatically included in the corresponding buffer so that it can be focused on in subsequent incremental learning.
[0271] Trigger a model update once a week, unlocking only the parameters of the last layer of the network for fine-tuning (the rest of the layers are frozen) to avoid catastrophic forgetting;
[0272] Under small sample conditions, the model can gradually adapt to extreme or novel scenarios by learning new samples in the buffer.
[0273] The specific results of comparing this system with a baseline system (traditional fusion strategy) on the extreme test set and the basic training set are shown in Table 1:
[0274] Evaluation Dimensions This system Baseline system Improvement First frame recognition accuracy 83.6% 67.2% +16.4% Steady-state error (mm) ±0.9 ±1.8 —50% False alarm decay rate (24h) 72.3% 45.1% +27.2% Consistency across materials 91.5% 78.4% +13.1% Anti-interference ability 88.2% 63.7% +24.5%
[0275] Table 1 Performance evaluation table of this embodiment and prior art
[0276] Based on Table 1, it can be found that the present system achieved an accuracy of 83.6%, significantly higher than the baseline system's 67.2%, indicating that it can more quickly identify ice and snow types in the initial detection stage. The root mean square error of ice thickness prediction was reduced from ±1.8mm to ±0.9mm, a reduction of half, demonstrating the high accuracy of the fusion algorithm. : After 24 hours of self-correction, the present system reduced the number of false alarms by 72.3%, while the baseline system only reduced it by 45.1%, indicating that the system has a stronger ability to correct its own predictions. The difference in detection results between the two materials (steel bridge vs. concrete bridge) was reduced by 13.1%, indicating that the algorithm has enhanced adaptability to differences in bridge deck materials. Interference such as headlight reflections and cracks are better suppressed, and the present system has improved by 24.5% compared to the baseline system.
[0277] With only 300 sets of training data, the system's first-frame recognition accuracy reached 83.6%, demonstrating its strong generalization ability in the data-scarce stage; the GWU (gated weighted update) mechanism enables the feature fusion weights to be automatically adjusted according to the scene (frost, ice, compacted snow, etc.), and compared with the fixed weight strategy, the accuracy is improved by 12.8%; after 4 weeks of progressive optimization (fine-tuning the last layer of the network every week), the extreme case recognition rate increased from the initial 71.5% to 89.3%, and the adaptive ability was significantly enhanced.
[0278] To verify the application effect of the system in a real scenario, we selected a steel box girder section of a cross-sea bridge in northern China and monitored the icing of the inversion layer at night. The specific process is as follows:
[0279] On-site conditions: Ambient temperature was approximately -3°C, wind speed was 6m / s, and ice was detected on the bridge deck at the inversion layer, with an initial ice thickness of 1.2mm.
[0280] System response: The first frame recognition took 1.8 seconds, indicating that ice formation was detected and a secondary warning was triggered.
[0281] Thickness prediction: The LSTM model predicts that the ice thickness will increase to 2.5 mm in 30 minutes, and the system will upgrade the warning to level 1.
[0282] The FAHP model output scheme priority is as follows:
[0283] Option B (liquid de-icing agent): comprehensive score 0.81
[0284] Solution C (microwave ice melting): comprehensive score 0.76
[0285] Solution A (mechanical deicing): comprehensive score 0.69
[0286] After selecting Option B (liquid de-icing agent), the ice layer completely melted within 20 minutes and the friction coefficient of the bridge surface returned to 0.48 (higher than the safety threshold of 0.35); compared with the traditional spraying amount, the amount of de-icing agent used was reduced by 24%, with better economic and environmental benefits.
[0287] This field validation demonstrates that the system is capable of industrial deployment even during the cold start phase, rapidly approaching optimal performance through mechanisms such as incremental learning and dynamic weighting. Furthermore, the system can continuously improve model accuracy and decision-making efficiency through self-learning in the event of more complex weather conditions.
[0288] In summary, the verification results of this embodiment show that: under the condition of limited initial data scale, the proposed bridge deck ice and snow warning system can still effectively complete ice and snow identification and thickness prediction, and continuously improve the recognition accuracy of extreme cases through means such as elastic memory pool and progressive fine-tuning, providing reliable technical support for road safety in northern winter.
[0289] Example 3
[0290] The present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the bridge surface ice and snow warning method based on the fusion of neural network and multimodal data as described in the above technical solution is implemented.
[0291] Example 4
[0292] The present invention provides an electronic device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the bridge surface ice and snow warning method based on the fusion of neural network and multimodal data as described in the above technical solution.
[0293] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0294] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0295] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0296] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0297] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.
[0298] The contents not described in detail in this specification belong to the prior art known to those skilled in the art.
Claims
1. A bridge surface ice and snow early warning method based on neural network and multimodal data fusion, characterized by: The following steps are involved: forming a multimodal feature representation based on an image of the bridge deck to be evaluated, meteorological parameters, and a bridge deck friction coefficient; the multimodal feature representation comprising a temperature-visual feature and an ice and snow texture feature; The multimodal feature representations are fused through a learnable dynamic weighting mechanism to obtain a unified feature embedding representation; Inputting the time series represented by the feature embedding into a time series neural network model to predict the future ice thickness of the bridge deck; Generate warning information based on predicted ice thickness.
2. The method according to claim 1, wherein: The following steps are also included: Calculating a normalized effectiveness value of each decision objective of a plurality of candidate de-icing schemes based on the predicted ice thickness, and calculating an evaluation value of each candidate de-icing scheme based on a corresponding weight of the decision objective to select an optimal de-icing scheme; The multiple decision-making objectives include at least safety, economy and environmental protection.
3. The method according to claim 1, wherein: The following steps are also included: The fused feature embedding representation is stored in a vector database to construct a historical ice and snow event feature library; a graph-based approximate nearest neighbor indexing algorithm is used to construct a topological structure for the feature library, thereby achieving fast nearest neighbor retrieval of the current feature embedding representation; the retrieved historical feature vectors and the time series of the feature embedding representation are jointly input into a temporal neural network model to provide a historical context reference for ice thickness prediction.
4. The method according to claim 1, wherein: When the predicted ice thickness or bridge friction coefficient is lower than the corresponding set threshold, an early warning message is output.
5. The method according to claim 1, wherein: The process of generating the temperature-visual feature includes: The multispectral image of the bridge deck and the corresponding temperature field matrix are spliced channel by channel to form a multi-channel input, which includes at least a visible light image data channel and a temperature data channel; The multi-channel image is first normalized and randomly masked, and then input into a pre-trained visual encoder based on a Transformer structure. The parameters of the front layer of the visual encoder are partially frozen to maintain the pre-trained visual representation, and the back layer introduces a temperature field attention mechanism to give higher weight to temperature-sensitive areas; After processing by the visual encoder, a high-dimensional feature vector of fused temperature and visual information is output for subsequent multimodal feature fusion processing.
6. The method according to claim 1, wherein: The process of generating the ice and snow texture features includes: A convolutional neural network-based backbone model is used to freeze some convolutional layers or residual blocks, and to fuse bridge deck image data, meteorological data, and bridge deck friction coefficient data into inputs. After preliminary feature extraction, a spatial pyramid pooling layer is added after the last residual block of the network. The spatial pyramid pooling layer uses multiple pooling windows of different sizes to pool the output features and concatenates the pooling results of each scale into a high-dimensional vector. After batch normalization and nonlinear activation function processing, the multidimensional ice and snow texture features are output.
7. The method according to claim 6, characterized in that The input data preprocessing steps of the backbone model include: The bridge deck image data, meteorological data, and bridge deck friction coefficient data are stored in a unified format in a database or file system. After reading the data from the database or file system, the bridge deck image data is subjected to data enhancement processing such as random rotation and color dithering. At the same time, the read meteorological data and bridge deck friction coefficient data are matched and jointly preprocessed with the processed image data in a tabular form, so that when input into the backbone model, all modal data are unified into a standardized format.
8. The method according to claim 1, characterized in that The process of fusing the multimodal feature representations through a learnable dynamic weighting mechanism includes: First, the cosine similarity between the temperature-visual features and the ice and snow texture features is calculated; When the similarity reaches or exceeds the preset threshold, the original weights are directly used to perform weighted fusion on the two features to obtain the fused features; When the similarity is lower than the preset threshold, the weight of the temperature-visual feature is reduced according to the difference between the cosine similarity and the threshold, and the weight of the ice and snow texture feature is increased accordingly. Finally, the adjusted weights are used to perform weighted fusion on the two features to obtain the final fusion features.
9. The method according to claim 1, wherein: The fuzzy analytic hierarchy process decision model constructs a fuzzy judgment matrix and uses the fuzzy eigenvalue method to calculate the comprehensive weight of each decision target.
10. A bridge surface ice and snow warning system based on neural network and multimodal data fusion, characterized by: For implementing the method according to any one of claims 1 to 9, comprising: A multimodal feature generation module, configured to generate a multimodal feature representation based on an image of the bridge deck to be evaluated, meteorological parameters, and a bridge deck friction coefficient; the multimodal feature representation includes temperature-visual features and ice and snow texture features; A multimodal feature fusion module, configured to fuse the multimodal feature representations through a learnable dynamic weighting mechanism to obtain a unified feature embedding representation; An ice thickness prediction module is used to input the time series represented by the feature embedding into a time series neural network model to predict the future ice thickness of the bridge deck; The warning generation module is used to generate warning information based on the predicted ice thickness. A decision generation module is configured to calculate a normalized effectiveness value for each decision objective of a plurality of candidate de-icing schemes based on the predicted ice thickness, and to calculate an evaluation value for each candidate de-icing scheme based on the corresponding weights of the decision objectives, so as to select an optimal de-icing scheme; the plurality of decision objectives including at least safety, economy, and environmental protection.
Citation Information
Patent Citations
Method and system for removing ice and snow on bridge floor based on neural network, terminal and medium
CN118940804A
Indoor ice surface defect detection method based on deep learning
CN119672292A
Method of providing companion animal angel book ashes platform business model that permanently preserves sense of remembrance and honor of companion animals by watching history of companion animals in video
KR102396956B1
Cited By
Multi-mode large model and light-weight small model collaborative road surface ice coagulation state prediction method
CN121524555A
Road icing emergency early warning system and method based on Internet of Things large model, and medium
CN122067421A
Road icing emergency early warning system and method based on internet of things large model and medium
CN122067421B