Radar, vision and Beidou integrated bridge collision monitoring method and system

Through the integrated radar, vision, and Beidou bridge collision monitoring methods, the limitations of a single sensor in complex environments and the heterogeneity of multimodal data is solved, high-precision and real-time collision risk monitoring are achieved, and bridge operation safety is ensured.

CN119961888AInactive Publication Date: 2025-05-09HUNAN XIANGYINHE SENSOR TECH CO LTD

Patent Information

Application Number
CN202510442941.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Due to the limitations of a single sensor in complex environments, existing bridge collision monitoring systems cannot conduct high-precision and real-time assessment of target collision risks. At the same time, the heterogeneity of multimodal data in physical properties, sampling rate and distribution characteristics makes it difficult to fusion data.

Method used

The bridge collision monitoring method of integrated radar, vision, and Beidou is adopted to collect data and preprocess it through radar, vision and Beidou modules. The information entropy is used to calculate the amount of information of the data and dynamically allocate weights. The data is aligned through the optimal transmission model, and combined with the deep learning network to integrate multimodal features to calculate the collision risk probability of the target, and trigger an alarm when the risk exceeds the threshold.

Benefits of technology

It realizes efficient fusion and alignment of multimodal data, eliminates the limitations of a single sensor, improves the accuracy and real-time nature of collision risk monitoring, and ensures comprehensive coverage and efficient identification of the surrounding environment of the bridge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961888A_ABST
    Figure CN119961888A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of bridge safety monitoring, and discloses a radar, vision and Beidou integrated bridge collision monitoring method and system, and the method comprises the following steps: S1, collecting the distance, speed and angle data of a target through a radar module, obtaining the real-time video data through a vision module, extracting the type and position information of the target, and carrying out the real-time video data; high-precision three-dimensional positioning data of a target is obtained through a Beidou module, and time synchronization of multi-modal data is realized based on a Beidou timestamp. Radar, vision and Beidou three-mode data are fused, and information entropy dynamic weight distribution and random optimal transmission alignment technologies are utilized, so that centimeter-level precision bridge collision monitoring is realized. The system has high real-time performance and strong environmental adaptability, and can dynamically adjust modal weight to adapt to a complex environment; meanwhile, multiple target types are covered, the reliability of risk assessment is improved, and comprehensive and efficient technical support is provided for bridge safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bridge safety monitoring, and in particular to a bridge collision monitoring method and system integrating radar, vision and Beidou. Background Art

[0002] As a key infrastructure in the transportation system, the safety of bridges is of great significance to the national economy and social development. With the increasing number and scale of modern bridges, the complexity of the environment around bridges is also increasing. For example, over-height vehicles commonly seen on highway bridges, ship collisions around river bridges, and floating objects in the natural environment may pose a threat to bridge structures. Therefore, establishing an efficient bridge collision monitoring system is an important means to ensure the safety of bridge operations.

[0003] Traditional bridge collision monitoring systems are usually based on the application of a single sensor, such as radar, vision or Beidou system. Radar can provide high-precision target distance and speed information, suitable for all-weather dynamic monitoring; the vision system has rich spatial information collection capabilities, can identify target types and provide intuitive video images; Beidou system is widely used in real-time tracking of targets with its high-precision positioning and time synchronization characteristics. However, the performance of a single sensor in a complex environment is often limited. For example, radar has insufficient target classification capabilities, the vision system is easily affected by light and weather, and Beidou's positioning performance may decline in areas with severe occlusion.

[0004] Multimodal sensor fusion technology provides a reliable solution for bridge collision monitoring. By fusing radar, vision and BeiDou system data, it can combine the advantages of different sensors to achieve comprehensive analysis and accurate prediction of target dynamic behavior. However, the heterogeneity of multimodal data needs to be solved in the data fusion process, such as differences in physical properties, distribution characteristics and sampling rates. At the same time, the quality of different sensor data may vary significantly in complex environments, so it is necessary to dynamically adjust the fusion strategy to ensure the real-time and accuracy of the monitoring system.

[0005] In recent years, the rapid development of deep learning and optimal transmission theory has provided new technical means for the efficient fusion of multimodal data. Target detection and feature extraction technology based on deep learning can extract multidimensional features from complex data, while optimal transmission theory performs well in data alignment by optimizing the matching relationship between distributions. The combination of these technologies has significantly improved the real-time, accuracy and robustness of multimodal data fusion, providing important technical support for the further development of bridge collision monitoring systems. Summary of the invention

[0006] In view of the shortcomings of the existing technology, the present invention provides a bridge collision monitoring method and system integrating radar, vision and Beidou, which solves the problem that in the dynamic target monitoring around the bridge, the target collision risk cannot be evaluated with high precision and in real time due to the limitations of a single sensor in a complex environment; at the same time, it solves the problem that the heterogeneity of multimodal data in physical properties, sampling rate and distribution characteristics makes data difficult to fuse.

[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions: a bridge collision monitoring method integrating radar, vision and Beidou, comprising the following steps:

[0008] S1. Collect the distance, speed and angle data of the target through the radar module, extract the target category and location information through the visual module to obtain real-time video data, obtain the high-precision three-dimensional positioning data of the target through the Beidou module, and realize the time synchronization of multi-modal data based on the Beidou timestamp;

[0009] S2. Preprocess the collected multimodal data, including filtering, denoising and feature extraction of radar data, target detection and feature extraction of visual data, and standardization of location information of Beidou data;

[0010] S3. Calculate the amount of information of each modal data based on information entropy, and dynamically assign weights of modal data according to the size of information entropy;

[0011] S4. Align the data of different modes through the optimal transmission model, and map the radar, vision and Beidou data into a unified feature space;

[0012] S5. Input the aligned multimodal features into the fusion model, use the deep learning network to fuse radar, vision and Beidou features, and calculate the collision risk probability of the target;

[0013] S6. When the collision risk probability exceeds the preset threshold, an alarm is triggered and the time, location and target category of the collision event are output.

[0014] Preferably, in step S1, the radar module calculates the distance, speed and angle of the target through a signal processing method, and obtains time series data for describing the motion characteristics of the target.

[0015] Preferably, in step S2, the visual module extracts the category and position information of the target through a target detection algorithm, and extracts the motion trajectory characteristics of the target using image processing technology.

[0016] Preferably, in the step S3, the weight distribution ratio of the data is determined by dynamically evaluating the amount of information of each modal data, and the weight distribution ratio is proportional to the amount of information of each modal data.

[0017] Preferably, in the step S4, the distribution of data of different modalities is aligned through an alignment algorithm based on transmission cost optimization, and the radar, vision and Beidou data are uniformly mapped to the same feature space.

[0018] Preferably, in the step S5, the fusion model extracts important features of each modal data based on a multi-head attention mechanism, and fuses the feature information of radar, vision and Beidou through a weight adjustment mechanism.

[0019] Preferably, in step S5, the collision risk probability is predicted by a deep learning model, the model input is the fused multimodal features, and the output is the probability value of the collision risk.

[0020] Preferably, in step S6, the alarm module includes a threshold trigger mechanism and an information recording module, and when the alarm is triggered, the time, location and target category of the collision event are recorded.

[0021] A bridge collision monitoring system integrating radar, vision and Beidou, including:

[0022] Radar module, used to obtain the distance, speed and angle information of the target;

[0023] A vision module that acquires real-time video around the bridge and extracts object categories and location information;

[0024] Beidou module, used to obtain high-precision three-dimensional position information of the target and provide time synchronization;

[0025] The data processing module includes a data preprocessing submodule, an information volume calculation submodule, and a data alignment submodule, which are used to standardize, dynamically assign weights, and align distribution of multimodal data;

[0026] Fusion and risk assessment module, including deep fusion model and risk analysis submodule, used for multi-modal data fusion and risk assessment of collision events;

[0027] The alarm module is used to trigger an alarm signal when the collision risk exceeds a preset threshold and output the time, location and category information of the collision event.

[0028] Preferably, the information amount calculation submodule in the data processing module performs weight allocation on radar, vision and Beidou data based on a dynamic evaluation method of multimodal data to improve the accuracy of the fusion result.

[0029] The present invention provides a bridge collision monitoring method and system integrating radar, vision and Beidou. It has the following beneficial effects:

[0030] 1. The present invention achieves efficient fusion and alignment of multimodal data by integrating data from three sensors: radar, vision and Beidou, combined with information entropy dynamic weight allocation and random optimal transmission alignment technology, effectively eliminating the limitations of a single sensor. Radar provides accurate target distance and speed information, vision provides target category and spatial characteristics, and Beidou provides high-precision three-dimensional position positioning. The three complement each other, so that the monitoring accuracy of collision risk reaches centimeter level, providing accurate monitoring and alarm for potential threats such as ships and vehicles around the bridge.

[0031] 2. The present invention accelerates the calculation speed of random optimal transmission by using a lightweight deep learning model and the Sinkhorn-Knopp optimization algorithm, ensuring that the system can process multimodal data in real time. At the same time, the system has the ability to dynamically adjust weights and can adapt to data changes in various complex environments (such as low light, rain, fog, and occlusion). When the visual module is affected by illumination or occlusion, the system can increase the weight of radar and Beidou data, thereby ensuring the reliability and real-time monitoring in complex environments.

[0032] 3. The present invention calculates the amount of information of each modal data through variational information entropy and dynamically adjusts the weight, thus realizing a flexible weight allocation mechanism for different scenarios and data quality. The contribution of radar, vision and Beidou is dynamically adjusted as environmental conditions change, thereby ensuring the robustness and flexibility of the system. Even if a certain modality has a low amount of information due to interference from the external environment, the system can still ensure the accuracy of collision risk assessment through the supplement of other modalities.

[0033] 4. The multimodal monitoring system of the present invention achieves comprehensive coverage of the environment around the bridge. Whether it is a ship in the water, a vehicle on land, or a floating object, the system can efficiently identify and predict collision risks. By comprehensively analyzing the time series, spatial position, and dynamic behavior of different targets, it can warn of potential collision risks and provide strong technical support for bridge operation safety. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 It is a schematic diagram of the steps of the bridge collision monitoring method of the present invention;

[0035] Figure 2 It is a schematic diagram of the bridge collision monitoring system module in the present invention. DETAILED DESCRIPTION

[0036] The following will be combined with the drawings in the specification of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0037] Example:

[0038] Please see attached Figure 1 The embodiment of the present invention provides a bridge collision monitoring method and system integrating radar, vision and Beidou, including the following steps:

[0039] S1. Collect the distance, speed and angle data of the target through the radar module, extract the target category and location information through the visual module to obtain real-time video data, obtain the high-precision three-dimensional positioning data of the target through the Beidou module, and realize the time synchronization of multi-modal data based on the Beidou timestamp;

[0040] S2. Preprocess the collected multimodal data, including filtering, denoising and feature extraction of radar data, target detection and feature extraction of visual data, and standardization of location information of Beidou data;

[0041] The data preprocessing step is the basis for realizing the fusion of radar, vision and Beidou trimodal data, and aims to provide standardized, time-series and characterization data input for subsequent information calculation, data alignment and risk assessment. This step ensures the accuracy and consistency of data processing by denoising and extracting dynamic features of radar signals, detecting targets and extracting spatial features of visual data, and processing time synchronization and position information of Beidou data. It should be noted that all processing operations in data preprocessing provide a unified feature space and timestamp for subsequent steps, which is crucial for realizing multimodal fusion.

[0042] Preprocessing of radar data:

[0043] In this embodiment, the radar module is used to collect the distance, speed and angle information of the target, which mainly includes denoising processing, signal analysis and feature extraction.

[0044] Specifically, the original signal output by the radar module is:

[0045] ;

[0046] in is the signal amplitude, is the carrier frequency, In order to calculate the motion characteristics of the target, the signal needs to be analyzed.

[0047] Distance to target By formula:

[0048] ;

[0049] Calculate, where c is the speed of light, B is the radar bandwidth, For frequency shift.

[0050] The velocity v of the target is determined by the Doppler effect using the following formula:

[0051]

[0052] in is the phase change, For the time interval.

[0053] As an option, in order to reduce the interference of noise on the signal, in this embodiment, the radar signal is low-pass filtered to eliminate invalid echoes and environmental noise.

[0054] It should be noted that when extracting the dynamic features of the target, a long short-term memory network (LSTM) is used to process the time series signal of the radar. For example, the recursive formula of LSTM is as follows:

[0055]

[0056] in For time series input, is the hidden layer output, , , b is the weight and bias of the network.

[0057] It can be understood that the use of LSTM is able to capture the dynamic changes of the target motion trajectory and provide reliable time series features for subsequent collision risk assessment.

[0058] Preprocessing of visual data:

[0059] In this embodiment, the visual module is used to extract the category, location and trajectory features of the targets around the bridge. After the original video data is processed by the deep learning model, the category label and spatial features of the target are output.

[0060] Specifically, target detection uses the YOLO model, which is achieved by optimizing the following objective function:

[0061]

[0062] in:

[0063] is the target classification loss, used for classification tasks.

[0064] is the regression loss of the bounding box, used to locate the target.

[0065] As an option, the extracted visual features are further processed by a convolutional neural network (CNN) to obtain the spatial distribution and motion trajectory of the target. For example, the features extracted by CNN are used to calibrate the target position and track its dynamic behavior.

[0066] It should be noted that in order to ensure the synchronization of visual data with radar and Beidou data, the time series of the visual module is aligned based on the Beidou timestamp in this embodiment to avoid feature fusion errors caused by time deviation.

[0067] Preprocessing of Beidou data:

[0068] In this embodiment, the Beidou module is used to provide high-precision three-dimensional position information of the target and serve as a reference for time synchronization. The data output by Beidou includes longitude, latitude and altitude information.

[0069] Specifically, the time synchronization formula of the Beidou module is:

[0070]

[0071] in , , They are the data timestamps of radar, vision and BeiDou respectively.

[0072] It should be noted that the Beidou data is standardized in this embodiment to unify the spatial distribution of the three-modal data. For example, by dynamically updating the Beidou data, centimeter-level spatial precision positioning can be achieved, providing support for accurate assessment of subsequent collision risks.

[0073] S3. Calculate the amount of information of each modal data based on information entropy, and dynamically assign weights of modal data according to the size of information entropy;

[0074] The information calculation and weight distribution steps are used to evaluate the importance of radar, vision and Beidou tri-modal data, and dynamically adjust the fusion weights of each modal data to adapt to different environments and changes in data quality. The characteristic information of each modality is calculated by variational information entropy to ensure that the weight distribution is consistent with the data quality and information contribution, thereby improving the accuracy and robustness of multi-modal data fusion. It should be noted that this step lays a theoretical foundation for subsequent data alignment and deep fusion, and ensures the adaptability and reliability of the system in complex scenarios.

[0075] Information volume calculation:

[0076] In this embodiment, the calculation of information volume is based on information entropy theory. Represents a random variable The uncertainty of is defined as:

[0077]

[0078] in is modal data The probability density function of .

[0079] As an option, in order to simplify the entropy calculation of high-dimensional data, the variational inference method is used in this embodiment. Approximate the posterior distribution of modal data Specifically, the variational expression of information entropy is:

[0080]

[0081] in:

[0082] is a variational distribution, usually a Gaussian distribution To ensure the efficiency of numerical calculations.

[0083] is a joint distribution, which represents the generative characteristics of the modal data x.

[0084] It should be noted that the calculation result of variational information entropy is used to measure the contribution of each modal data. The radar, vision and Beidou modal data are respectively denoted as , , , the corresponding entropy value calculation formula is:

[0085]

[0086]

[0087]

[0088] It should be noted that the calculation steps of the above formula include joint distribution modeling, sampling and numerical integration of each modal data to ensure that the information entropy value can accurately reflect the uncertainty and information content of the data.

[0089] Weight distribution:

[0090] In one possible implementation, based on the calculation result of information entropy, a proportional allocation mechanism is used to dynamically adjust the fusion weight of each modality. Specifically, the weights of radar, vision, and Beidou modalities are , , Calculated by the following formulas:

[0091]

[0092] It should be noted that the above-mentioned weight distribution mechanism ensures that modal data with large amounts of information have a higher influence in the fusion process, thereby enhancing the robustness of the system.

[0093] As an option, the environmental adaptability of modal data is also considered in this embodiment. For example, when radar data is disturbed by rain and fog or visual data is degraded due to insufficient light, the information entropy value will decrease accordingly, thereby reducing its proportion in the weight distribution. This dynamic weight adjustment mechanism effectively reduces the impact of noise and invalid data on the multimodal fusion results.

[0094] It can be understood that the dynamic allocation of weights provides an optimized reference basis for subsequent modal data alignment and deep fusion, so that the present invention has higher reliability in complex environments.

[0095] S4. Align the data of different modes through the optimal transmission model, and map the radar, vision and Beidou data into a unified feature space;

[0096] The data alignment step is used to solve the heterogeneity of radar, vision and Beidou tri-modal data in terms of physical properties, distribution and sampling rate. The data distribution of different modes is nonlinearly aligned through the random optimal transmission algorithm to ensure the consistency of multi-modal data expression in the same feature space. It should be noted that this step provides high-quality input data for the subsequent deep fusion model, significantly improving the accuracy and stability of multi-modal fusion.

[0097] Data alignment:

[0098] In this embodiment, in order to achieve data alignment, it is first necessary to define the distribution of different modal data. Specifically, the distribution of radar, vision and Beidou modal data is respectively recorded as , and Represents the distribution density of each modal data in the feature space.

[0099] In one possible implementation, random optimal transmission theory is used to align modal data. Random optimal transmission introduces the uncertainty of noise perturbation modeling distribution on the basis of traditional optimal transmission, and the goal is to minimize the transmission cost by optimizing the joint distribution between modes.

[0100] The goal of multimodal data alignment is to solve the heterogeneity problem between different data sources through mathematical models, including differences in the distribution of modal data, differences in sampling rates, and differences in spatial coordinates. To achieve this goal, the optimal transport (OT) theory is used, which aims to find a way to match different modal data and minimize the transmission cost between data distributions.

[0101] Assume that the distributions of radar data, visual data, and BeiDou data are , and , hoping to map the features of these three modes into a unified feature space through the optimal transmission method. The goal of the optimal transmission model is to find a joint distribution To minimize the "transmission cost" between each modal data.

[0102] As an option, the transmission cost function Indicates the difference between modal data. Specifically, the transmission cost It is usually chosen to be the square of the Euclidean distance between the modal data. Assume and are data points from different modalities, then the transmission cost is defined as:

[0103]

[0104] Where x and y represent the feature points of the two modes respectively. represents the Euclidean distance.

[0105] The core of optimal transmission is to find an optimal joint distribution by minimizing the objective function. The mathematical expression of the optimal transmission problem is:

[0106]

[0107] in is the joint distribution of modal data, Represents modal data and All possible joint distributions of .

[0108] In order to further consider the uncertainty in the environment, such as the inconsistency of modal data caused by factors such as weather and occlusion, this embodiment introduces random variables into the optimal transmission model. , this uncertainty is resolved through random optimal transmission. The goals of random optimal transmission are:

[0109]

[0110] in, is a normally distributed noise variable used to simulate disturbances in dynamic scenes.

[0111] Represents a Gaussian distribution of noise, used to model the uncertainty of the modal data.

[0112] In order to efficiently solve the above random optimal transmission problem, this embodiment uses the Sinkhorn-Knopp algorithm for numerical optimization. Exemplarily, the Sinkhorn-Knopp algorithm transforms the optimization problem into a solvable form by introducing a regularization term, and its objective function is expressed as:

[0113]

[0114] in:

[0115] is the regularization parameter used to control the stability of the calculation;

[0116] represents the joint probability between feature points of radar and visual data.

[0117] Specifically, in actual implementation, the Sinkhorn-Knopp algorithm iteratively updates The value of gradually approaches the optimal joint distribution. It can be understood that the introduction of this algorithm effectively improves the computational efficiency of random optimal transmission and ensures the real-time requirements.

[0118] Feature Mapping:

[0119] It should be noted that after alignment is completed by the above random optimal transmission method, radar, vision and Beidou data are mapped into the same feature space. Exemplarily, the mapped features can be uniformly expressed as:

[0120]

[0121] in , , Represent the mapping functions of radar, vision and Beidou modal features respectively.

[0122] In one possible implementation, the feature map can be further combined with dimensionality reduction techniques such as principal component analysis (PCA) to reduce the redundancy of the feature space and improve the efficiency of subsequent deep fusion.

[0123] S5. Input the aligned multimodal features into the fusion model, use the deep learning network to fuse radar, vision and Beidou features, and calculate the collision risk probability of the target;

[0124] The multimodal fusion and collision risk assessment step is a key link based on the fusion of radar, vision and Beidou data. Through the combination of weighted fusion and deep learning models, the features of heterogeneous modalities are integrated into a unified representation, and the probability prediction and position assessment of collision risk are completed. It should be noted that this step not only relies on the aforementioned data preprocessing and alignment results, but also fully considers the importance of each modal feature, optimizes the fusion effect through dynamic weight allocation, and provides solid technical support for high-precision collision risk assessment.

[0125] Multimodal fusion:

[0126] In this embodiment, multi-modal fusion is accomplished by a weighted fusion method. Specifically, the aligned radar features , visual features and BeiDou features According to their respective weights , , After weighted integration, the fused features are expressed as:

[0127]

[0128] in:

[0129] , , is the weight dynamically calculated according to information entropy.

[0130] , , Respectively, they are the representations of radar, vision and Beidou features in the unified feature space.

[0131] The random optimal transmission model is used to solve the problem of heterogeneous distribution of radar, vision and BeiDou data. Specifically, it is assumed that radar, vision and BeiDou data obey different probability distributions. , and , the goal of the optimal transmission problem is to find an optimal joint distribution To minimize the transmission cost. The optimal transmission model solves the Wasserstein distance:

[0132]

[0133] in, represents all possible joint distributions between modal data, is the cost function, which measures the difference between different modal data.

[0134] As an option, the weighted fused feature F can be further processed by feature normalization to eliminate the impact of scale differences between different modal features on the deep learning model. Specifically, the normalization formula is:

[0135]

[0136] in and They represent the mean and standard deviation of the fusion feature F respectively.

[0137] It should be noted that the representation of the fusion feature F combines the time series, spatial position and motion dynamic information of radar, vision and Beidou modes, providing comprehensive input for subsequent collision risk assessment.

[0138] Collision risk assessment:

[0139] In this embodiment, the collision risk assessment is implemented through a deep learning model. Specifically, a Transformer-based network architecture is adopted to extract key information from fused features through a multi-head attention mechanism.

[0140] The goal of the data fusion model is to fuse the features of radar, vision and BeiDou by weighted fusion for subsequent collision risk assessment. In this embodiment, a multi-head self-attention mechanism is selected, which can capture the dependencies between different modalities when processing different modal features, thereby improving the prediction accuracy of the model.

[0141] The principle of self-attention mechanism:

[0142] The key idea of ​​the self-attention mechanism is to calculate the relationship between each position in the input data and then weight different features. Specifically, given the input feature matrix , the query, key and value matrices are obtained by linear transformation respectively:

[0143] Query Matrix ;

[0144] Key Matrix ;

[0145] Value Matrix .

[0146] The core calculation formula is:

[0147]

[0148] in:

[0149] Q, K, and V are query, key, and value matrices, respectively;

[0150] is the inner product of the query and the key;

[0151] is the normalization function;

[0152] The dimensions of the matrix used to scale the results to improve computational stability.

[0153] Multi-head attention mechanism:

[0154] In order to capture the multi-level dependencies between different features, the model uses a multi-head attention mechanism. By computing multiple attention heads in parallel and concatenating their outputs, the final attention representation is obtained. The output of the multi-head attention mechanism can be expressed as:

[0155]

[0156] in, It is The attention calculation of the head, is the output weight matrix, is the number of attention heads.

[0157] The weighted radar, vision and Beidou features are input into the multi-head self-attention mechanism for fusion. In this way, the model can dynamically adjust the contribution between different modalities, thereby achieving effective integration of multimodal information. Finally, the fused features It will be used for subsequent collision risk assessment.

[0158] As an option, the matrix input to the Transformer network is composed of the normalized fused features After calculation by the multi-head attention mechanism, high-order features for collision risk probability prediction are output.

[0159] In the risk probability calculation, this embodiment adopts a logistic regression model to complete the probability prediction of collision risk through the following formula:

[0160]

[0161] in:

[0162] W is the weight matrix of logistic regression;

[0163] b is the bias term;

[0164] is the Sigmoid activation function, defined as .

[0165] It should be noted that the output of logistic regression is represents the probability of target collision. For example, when Exceeding the set threshold When the vehicle is in a collision situation, the system will determine that there is a risk of collision and trigger an alarm.

[0166] S6. When the collision risk probability exceeds the preset threshold, an alarm is triggered and the time, location and target category of the collision event are output.

[0167] Please see attached Figure 2 As part of the present invention, the present invention also provides an embodiment, a bridge collision monitoring system integrating radar, vision and Beidou, comprising:

[0168] Radar module, used to obtain the distance, speed and angle information of the target;

[0169] Specifically:

[0170] The original signal output by the radar module is transformed by fast Fourier transform Processing, extract the distance d and speed v of the target. The calculation formula of distance d is:

[0171]

[0172] Where c is the speed of light, B is the radar bandwidth, For frequency shift.

[0173] The target's velocity v is calculated by the Doppler effect, using the formula:

[0174]

[0175] in, is the phase change, is the carrier frequency, For the time interval.

[0176] As an option, the radar module embeds a low-pass filter to remove environmental noise and interference signals. In addition, the dynamic trajectory information of the target is modeled through time series feature extraction methods (such as LSTM) to provide input for subsequent fusion.

[0177] A vision module that acquires real-time video around the bridge and extracts object categories and location information;

[0178] In this embodiment, the vision module includes a camera and an image processing unit, which are used to obtain real-time video data around the bridge and extract the category and location information of the target.

[0179] Specifically:

[0180] The image processing unit uses a deep learning target detection model (such as YOLOv8) to detect targets in the video stream. The category and location information of the target are obtained by optimizing the following objective function:

[0181]

[0182] in:

[0183] is the classification loss, which indicates the recognition error of the target category;

[0184] is the bounding box regression loss, which represents the prediction error of the target location.

[0185] The extracted visual features include the spatial distribution of the target, the bounding box position and the motion trajectory. As a possible implementation method, convolutional neural network (CNN) is used to further extract high-dimensional features to ensure the integrity of feature expression.

[0186] It should be noted that the visual module is synchronized through the timestamp provided by the Beidou module to ensure the consistency of the visual data with the radar and Beidou data on the timeline.

[0187] Beidou module, used to obtain high-precision three-dimensional position information of the target and provide time synchronization;

[0188] The Beidou module obtains high-precision three-dimensional position information of the target through real-time kinematic positioning (RTK) technology and provides time synchronization function.

[0189] Specifically:

[0190] The location information output by the Beidou module includes the longitude, latitude and altitude of the target , , In order to keep consistent with other modal data, the Beidou module standardizes the timestamp. The time synchronization formula is:

[0191]

[0192] in, , , Represents the timestamps of radar, vision and BeiDou data respectively.

[0193] It should be noted that the Beidou module has high environmental adaptability and can provide centimeter-level spatial accuracy in complex scenarios, providing important location information for the system's collision assessment.

[0194] The data processing module includes a data preprocessing submodule, an information volume calculation submodule, and a data alignment submodule, which are used to standardize, dynamically assign weights, and align distribution of multimodal data;

[0195] The data processing module includes a data preprocessing submodule, an information volume calculation submodule and a data alignment submodule, which are respectively used for standardization of raw data, dynamic evaluation of information volume and distribution alignment of modal data.

[0196] Specifically:

[0197] Data preprocessing submodule:

[0198] Denoising and time series modeling for radar data.

[0199] Perform object detection and feature extraction on visual data.

[0200] Standardize Beidou data and complete time synchronization.

[0201] Information volume calculation submodule:

[0202] The variational information entropy method is used to dynamically evaluate the importance of radar, vision and Beidou data. The calculation formula of information entropy is:

[0203]

[0204] Data alignment submodule:

[0205] The distribution of modal data is aligned based on random optimal transmission theory, and the optimization goal is:

[0206]

[0207] The alignment problem is efficiently solved using the Sinkhorn-Knopp algorithm.

[0208] Fusion and risk assessment module, including deep fusion model and risk analysis submodule, used for multi-modal data fusion and risk assessment of collision events;

[0209] The fusion and risk assessment module includes the deep fusion model and collision risk analysis sub-modules.

[0210] Specifically:

[0211] Deep Fusion Model:

[0212] According to the information calculation results, the weights of radar, vision and Beidou features are allocated, and the fusion features are expressed as:

[0213]

[0214] The key information in the fusion features is extracted through the Transformer-based network architecture. The calculation formula of the attention mechanism is:

[0215]

[0216] Collision risk analysis submodule:

[0217] The core of the collision risk assessment model is to process the fused features through a deep learning network and finally output the collision risk probability of the target. In order to perform probability prediction, this embodiment adopts a logistic regression model, and its output is the probability of a collision event.

[0218] The logistic regression model is used for binary classification problems and outputs the probability of the target event occurring. , the risk probability of a collision event is calculated by the logistic regression model, and the formula is:

[0219]

[0220] Where F is the fused feature vector, W is the weight matrix, and b is the bias term. is the Sigmoid activation function.

[0221] An alarm module is used to trigger an alarm signal when the collision risk exceeds a preset threshold and output the time, location and category information of the collision event;

[0222] The alarm module is used to trigger an alarm when the collision risk exceeds a preset threshold.

[0223] Specifically:

[0224] The alarm module receives the collision risk probability As a result, when (threshold), an alarm signal is triggered.

[0225] The output alarm information includes the time, location and target category of the collision event.

[0226] It should be noted that the alarm module supports wireless communication function and can send collision warning information to the bridge management system in real time.

[0227] In some embodiments, the system can be combined with additional sensors (such as infrared cameras) to enhance monitoring capabilities in bad weather. As a possible improvement, the data processing module can introduce adaptive filtering technology to further improve the adaptability of modal data to environmental changes.

[0228] The above description is the main implementation mode of the system of the present invention. Those skilled in the art can optimize and adjust the module design and data processing method based on the disclosed content to meet the needs of different application scenarios.

[0229] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A bridge collision monitoring method integrating radar, vision and Beidou, characterized in that: The following steps are involved: S1. Collect the distance, speed and angle data of the target through the radar module, extract the target category and location information through the visual module to obtain real-time video data, obtain the high-precision three-dimensional positioning data of the target through the Beidou module, and realize the time synchronization of multi-modal data based on the Beidou timestamp; S2. Preprocess the collected multimodal data, including filtering, denoising and feature extraction of radar data, target detection and feature extraction of visual data, and standardization of location information of Beidou data; S3. Calculate the amount of information of each modal data based on information entropy, and dynamically assign weights of modal data according to the size of information entropy; S4. Align the data of different modes through the optimal transmission model, and map the radar, vision and Beidou data into a unified feature space; S5. Input the aligned multimodal features into the fusion model, use the deep learning network to fuse radar, vision and Beidou features, and calculate the collision risk probability of the target; S6. When the collision risk probability exceeds the preset threshold, an alarm is triggered and the time, location and target category of the collision event are output.

2. The bridge collision monitoring method integrating radar, vision and Beidou according to claim 1 is characterized in that: In the step S1, the radar module calculates the distance, speed and angle of the target through a signal processing method, and obtains time series data for describing the motion characteristics of the target.

3. The bridge collision monitoring method integrating radar, vision and Beidou according to claim 1 is characterized in that: In step S2, the visual module extracts the category and location information of the target through the target detection algorithm, and extracts the motion trajectory characteristics of the target using image processing technology.

4. The bridge collision monitoring method integrating radar, vision and Beidou according to claim 1 is characterized in that: In the step S3, the weight distribution ratio of the data is determined by dynamically evaluating the amount of information of each modal data, and the weight distribution ratio is proportional to the amount of information of each modal data.

5. The bridge collision monitoring method integrating radar, vision and Beidou according to claim 1 is characterized in that: In the step S4, the distribution of data of different modalities is aligned through an alignment algorithm based on transmission cost optimization, and the radar, vision and Beidou data are uniformly mapped to the same feature space.

6. The bridge collision monitoring method integrating radar, vision and Beidou according to claim 1 is characterized in that: In the step S5, the fusion model extracts important features of each modal data based on a multi-head attention mechanism, and fuses the feature information of radar, vision and Beidou through a weight adjustment mechanism.

7. The bridge collision monitoring method integrating radar, vision and Beidou according to claim 1 is characterized in that: In step S5, the collision risk probability is predicted by a deep learning model, the model input is the fused multimodal features, and the output is the probability value of the collision risk.

8. The bridge collision monitoring method integrating radar, vision and Beidou according to claim 1 is characterized in that: In the step S6, the alarm module includes a threshold trigger mechanism and an information recording module, and when the alarm is triggered, the time, location and target category of the collision event are recorded.

9. A bridge collision monitoring system integrating radar, vision and Beidou, characterized in that: A bridge collision monitoring method integrating radar, vision and Beidou as described in any one of claims 1 to 8 comprises: Radar module, used to obtain the distance, speed and angle information of the target; A vision module that acquires real-time video around the bridge and extracts object categories and location information; Beidou module, used to obtain high-precision three-dimensional position information of the target and provide time synchronization; The data processing module includes a data preprocessing submodule, an information volume calculation submodule, and a data alignment submodule, which are used to standardize, dynamically assign weights, and align distribution of multimodal data; Fusion and risk assessment module, including deep fusion model and risk analysis submodule, used for multi-modal data fusion and risk assessment of collision events; The alarm module is used to trigger an alarm signal when the collision risk exceeds a preset threshold and output the time, location and category information of the collision event.

10. The bridge collision monitoring system integrating radar, vision and Beidou according to claim 9 is characterized in that: The information volume calculation submodule in the data processing module performs weight allocation on radar, vision and Beidou data based on a dynamic evaluation method of multimodal data to improve the accuracy of the fusion result.

Citation Information

Patent Citations

  • Bridge ship collision prevention monitoring and early warning system

    CN114333424A

  • Internal medicine nursing infection control and protection system

    CN118230917A

  • Lateral vehicle early warning system

    CN119218103A

  • Ambient sound event detection method based on multi-modal data fusion

    CN119446154A

Cited By

  • Self-adaptive cross validation thundersight data fusion method and system

    CN120508999A

  • Composite radar photoelectric detection system

    CN120522690A