Bridge state intelligent detection and evaluation system based on machine vision
By combining multi-modal data alignment and cross-modal feature fusion with confidence decision downgrading and visual linkage assessment, the problems of high computational complexity, long real-time response delay and insufficient environmental adaptability of existing bridge inspection systems are solved, and efficient and stable bridge condition detection is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG TRANSPORTATION INST
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-21
AI Technical Summary
Existing bridge health monitoring systems suffer from high computational complexity, long real-time response delays, insufficient environmental adaptability, and limited generalization capabilities. They struggle to effectively identify crack morphologies in rare or marginal cases, and multimodal data fusion further increases system integration complexity and cost.
The modal vibration signals and acoustic reflection time delay spectrum signals of the bridge structure are synchronously acquired by the multi-mode data alignment module. The structural state correlation index is extracted by the cross-modal dynamic correlation feature fusion model. Combined with the confidence decision degradation module and the visual linkage evaluation module, real-time monitoring and robust degradation are achieved, and the region of interest and segmentation weight of the visual detection model are dynamically adjusted.
While ensuring high detection accuracy, the real-time response latency was shortened, the system's computational efficiency and edge scene recognition capabilities were improved, the continuity and stability of detection were ensured, and the complexity of system integration was reduced.
Smart Images

Figure CN121904584A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bridge condition detection technology, specifically to a machine vision-based intelligent bridge condition detection and evaluation system. Background Technology
[0002] The intelligent bridge condition detection and evaluation system can automatically detect and evaluate the structural condition of bridges. Through intelligent algorithms, the system can identify various potential defects in bridges, including cracks, corrosion, and deformation, and scientifically assess the safety and performance of the bridges.
[0003] The existing technology, with publication number CN119888496A, entitled "A Bridge Health Status Detection and Early Warning Method and System Based on Image Recognition," relates to the field of bridge health monitoring. It includes: acquiring bridge crack images; preprocessing the bridge crack images to form a bridge crack image dataset; manually augmenting the bridge crack image dataset using a sliding window algorithm and extracting features to generate a feature dataset; training a model using a convolutional neural network to obtain a crack recognition model; collecting bridge health status data in real time using sensors and using the crack recognition model to detect cracks; modeling the bridge health status using time series analysis to generate dynamic damage prediction results; implementing bridge health status assessment and providing early warning of bridge damage; and improving the accuracy and robustness of crack recognition by augmenting and extracting features from bridge crack images using an adaptive sliding window algorithm, combined with a convolutional neural network and a dynamic attention mechanism.
[0004] However, the above technical solutions have significant performance bottlenecks and applicability limitations. From a performance perspective, although multi-resolution feature fusion and dynamic attention mechanisms improve the model's recognition ability, the computational complexity is high, resulting in long real-time response delays, which makes it difficult to meet the needs of large-scale real-time monitoring. In addition, the model is highly dependent on the distribution of training data, has insufficient generalization ability, and is difficult to effectively identify crack morphology in rare or marginal cases. In terms of scene adaptability, existing solutions suffer from significant image quality degradation in extreme environments such as low light, rain, and snow, leading to a substantial decrease in crack detection accuracy. Furthermore, single visual sensors struggle to handle complex lighting and occlusion issues. While multimodal data fusion improves detection stability, it increases sensor cost and system integration complexity, posing challenges for practical deployment. Overall, current technologies, while ensuring high accuracy, suffer from insufficient response speed and environmental adaptability, and have limited ability to identify edge cases, creating a critical technological gap that urgently needs to be addressed in the field of intelligent bridge inspection.
[0005] The information disclosed in the background section above is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] The purpose of this invention is to provide a machine vision-based intelligent bridge condition detection and evaluation system to solve the problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention provides the following technical solution: A machine vision-based intelligent bridge condition detection and evaluation system includes: The multi-mode data alignment module is used to synchronously acquire the modal vibration signals and acoustic wave reflection time delay spectrum signals of the bridge structure. Through a high-precision timestamp synchronization mechanism and a dynamic time warping algorithm, the modal vibration signals and acoustic wave reflection time delay spectrum signals are time-series aligned and registered within a unified time window to generate a multi-source heterogeneous time series dataset. The correlation index generation module inputs the multi-source heterogeneous time-series dataset into a preset cross-modal dynamic correlation feature fusion model; the cross-modal dynamic correlation feature fusion model is constructed based on a multi-task learning framework, and uses a recurrent neural network to extract vibration mode features and acoustic delay features respectively, and calculates the structural state correlation index; The confidence decision degradation module is used to monitor the input confidence of modal vibration signals and acoustic wave reflection time delay spectrum signals in real time, and trigger a robust degradation mechanism when either input confidence is lower than a preset input confidence threshold. This includes automatically adjusting the feature fusion weights in the cross-modal dynamic correlation feature fusion model to reduce the impact of low-confidence signals, or blocking the output of the structural state correlation index and generating a switching command to control the system to enter a single visual detection mode; otherwise, the structural state correlation index is output as effective guidance information. The visual linkage evaluation module, in response to the effective guidance information, dynamically adjusts the coordinates of the region of interest and the semantic segmentation weights of the visual detection model based on the structural state correlation index, and generates refined recognition results and evaluation reports; or in response to the switching command, it locks the region of interest and resets the segmentation weights, performs basic visual detection, and generates an evaluation report with maintenance prompts.
[0008] Furthermore, the timing alignment and registration performed by the multi-modal data alignment module includes: Data acquisition is triggered using a precise time protocol to obtain the original vibration waveform and sound waveform with nanosecond-level timestamps; a short-time Fourier transform is performed on the original vibration waveform to extract the energy spectral density feature sequence, and the offset feature sequence of the first arrival time of the echo is extracted from the sound waveform. The energy spectral density feature sequence and the offset feature sequence are then resampled to a unified target frequency.
[0009] Furthermore, the timing alignment and registration performed by the multi-modal data alignment module also includes: A constraint window based on diagonal bandwidth is constructed. Within the constraint window, the Euclidean distance matrix between the energy spectral density feature sequence and the offset feature sequence is calculated. The dynamic time warping algorithm is used to find the warping path with the minimum cumulative distance. A one-to-one mapping relationship between time points is established based on the warping path.
[0010] Furthermore, the correlation index generation module includes: Using the dual-channel long short-term memory network encoder in the cross-modal dynamic correlation feature fusion model, the modal vibration features and acoustic delay features in the multi-source heterogeneous time series dataset are encoded into vibration latent state vectors and acoustic latent state vectors, respectively; based on the interaction between the vibration latent state vectors and the acoustic latent state vectors at each time step, a collaborative gain weight is calculated to characterize the degree of temporal resonance between the two. The vibration latent state vector and the acoustic latent state vector are weighted and aggregated using the cooperative gain weight to generate the structural state correlation index, which represents the joint probability of bridge structural stiffness degradation and internal damage propagation within the current time window.
[0011] Furthermore, the correlation index generation module specifically includes: The cooperative interaction energy of the vibration hidden state vector and the sound wave hidden state vector at the current time step is calculated based on the hyperbolic tangent activation function; the cooperative interaction energy is normalized by the normalized exponential function to generate the cooperative gain weight with a value range between 0 and 1. The vibration latent state vector and the acoustic latent state vector are concatenated, and the concatenated vector is weighted and aggregated using the cooperative gain weight to obtain the cross-modal fusion feature. The cross-modal fusion feature is then dimensionality-reduced through a fully connected layer and mapped using an S-shaped growth curve function to finally generate the structural state correlation index.
[0012] Furthermore, the robust degradation decision executed by the confidence decision degradation module includes: The information entropy of the modal vibration signal and the acoustic wave reflection time delay spectrum signal is calculated as the input confidence. This includes calculating the normalized information entropy of the modal vibration signal and the normalized information entropy of the acoustic wave reflection time delay spectrum signal for the signal time window at the current moment, and taking the reciprocal of the normalized information entropy of the two as the vibration input confidence and the acoustic wave input confidence, respectively.
[0013] Furthermore, the confidence decision degradation module compares the data integrity and reliability of the modal vibration signal and the acoustic wave reflection time delay spectrum signal based on a preset input confidence threshold, and determines whether the system is in a state of partial sensor failure or complete physical sensor failure based on the comparison result, and triggers the corresponding soft degradation strategy or hard switching strategy accordingly; the input confidence threshold includes a vibration confidence benchmark threshold and an acoustic position confidence benchmark threshold, and the classification decision includes executing a soft degradation strategy and a hard switching strategy.
[0014] Furthermore, in the confidence decision degradation module, when the input confidence of any signal is lower than the preset input confidence threshold but the other signal is normal, the function degradation maintenance when some sensors fail will execute a soft degradation strategy: in the cross-modal dynamic correlation feature fusion model, the attention weight corresponding to the low confidence signal will be forcibly set to zero. When both the vibration input confidence and the acoustic input confidence are below the input confidence threshold, a hard switching strategy is executed for the detection mode safety transfer in the event of full sensor failure. This includes cutting off the data path of the structural state correlation index and generating a switching command containing the visual dominant mode identifier bit and sending it to the visual linkage evaluation module.
[0015] Furthermore, when the visual linkage evaluation module responds to the effective guidance information, it specifically performs the following linkage control: analyzes the numerical strength of the structural state correlation index, calculates the center coordinates and cropping size of the region of interest of the visual detection model on the input image through a preset spatial mapping function, and realizes optical focusing on the damaged area; establishes a positive correlation mapping relationship between the value of the structural state correlation index and the weight of the loss function in the image semantic segmentation algorithm, and dynamically increases the pixel-level classification weight of the crack category as the value of the structural state correlation index increases, so as to suppress the interference of background texture noise on the identification of fine cracks.
[0016] Furthermore, upon receiving the switching instruction, the visual linkage evaluation module immediately interrupts data communication with the cross-modal dynamic correlation feature fusion model and forcibly locks the current input state of the visual detection model; it resets and locks the processing range of the visual detection model to the full field of view of the camera, and stops the dynamic coordinate adjustment based on the structural state correlation index to ensure that the maximum detection range is covered in the absence of prior guidance information. The classification weights of each category in the image semantic segmentation algorithm are restored to the preset initial equilibrium value, eliminating the weight bias for crack categories previously introduced by the high value of the structural state correlation index, and preventing overfitting false alarms when there is no internal data support. Based on the full field of view and the initial equalization value, the visual inspection model is run independently to identify cracks in the acquired images and generate a bridge health assessment report that includes a sensor data stream missing status identifier and the current visual inspection results.
[0017] Compared with the prior art, the beneficial effects of the present invention are: This invention utilizes a dual-channel long short-term memory network encoder and hyperbolic tangent activation function to calculate the structural state correlation index as effective guiding information through a correlation index generation module. The visual linkage evaluation module responds to this information by using a preset spatial mapping function to calculate the center coordinates and cropping size of the region of interest (ROI) on the input image for the visual detection model, achieving optical focusing on the damaged area. This reduces computational redundancy in invalid background areas and shortens real-time response latency. Simultaneously, a positive correlation mapping relationship is established between the structural state correlation index value and the weight of the loss function in the image semantic segmentation algorithm. As the structural state correlation index value increases, the pixel-level classification weight of the crack category is dynamically increased to suppress the interference of background texture noise on the identification of subtle cracks. This effectively solves the problem of insufficient generalization ability of existing models when facing rare or marginal crack morphologies, improving the system's computational efficiency and edge scene recognition capability while ensuring high detection accuracy. This invention also utilizes a confidence decision degradation module to calculate the normalized information entropy of the modal vibration signal and the normalized information entropy of the acoustic wave reflection delay spectrum signal in real time, and takes the reciprocal as the input confidence level. Based on a preset input confidence level threshold, a hierarchical judgment is performed. When some sensors fail, a soft degradation strategy is executed, that is, in the cross-modal dynamic correlation feature fusion model, the attention weights corresponding to low confidence signals are forcibly set to zero to maintain functional operation. When all sensors fail or extreme environments cause the input confidence level to be lower than the input confidence level threshold, a hard switching strategy is executed. A switching command containing the visual dominant mode identifier is generated, and the visual linkage evaluation module is controlled to reset and lock the processing range to the full field of view of the camera. The classification weights of each category in the image semantic segmentation algorithm are restored to the preset initial equilibrium value. Thus, even in scenarios where the quality of multimodal data deteriorates due to low light, rain, snow, etc., the continuity and stability of detection can still be ensured through safe mode transfer. This effectively overcomes the risk of system paralysis caused by poor environmental adaptability or sensor failure in traditional multimodal systems and reduces the practical deployment challenges brought about by system integration complexity. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of the overall system application of the present invention; Figure 2 This is a schematic diagram of the overall system framework of the present invention; Figure 3 This is a schematic diagram of the process framework of the confidence decision downgrade module and the visual linkage evaluation module of the present invention. Detailed Implementation
[0019] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0020] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below. Example
[0021] Please see Figure 1 and Figure 2 This invention provides a machine vision-based intelligent bridge condition detection and evaluation system, the specific steps of which include: The multi-mode data alignment module synchronously acquires the modal vibration signals and acoustic wave reflection time delay spectrum signals of the bridge structure. Through a high-precision timestamp synchronization mechanism and a dynamic time warping algorithm, the modal vibration signals and acoustic wave reflection time delay spectrum signals are time-series aligned and registered within a unified time window to generate a multi-source heterogeneous time series dataset. The correlation index generation module inputs the multi-source heterogeneous time-series dataset into a preset cross-modal dynamic correlation feature fusion model; the cross-modal dynamic correlation feature fusion model is constructed based on a multi-task learning framework, and uses a recurrent neural network to extract vibration mode features and acoustic delay features respectively, and calculates the structural state correlation index; The confidence decision degradation module monitors the input confidence levels of modal vibration signals and acoustic wave reflection time delay spectrum signals in real time. When any input confidence level falls below a preset input confidence threshold, a robust degradation mechanism is triggered. This mechanism includes automatically adjusting the feature fusion weights in the cross-modal dynamic correlation feature fusion model to reduce the impact of low-confidence signals, or blocking the output of the structural state correlation index and generating a switching command to control the system into a single visual detection mode. Otherwise, the structural state correlation index is output as effective guidance information. The visual linkage evaluation module, in response to the effective guidance information, dynamically adjusts the coordinates of the region of interest and the semantic segmentation weights of the visual detection model based on the structural state correlation index, and generates refined recognition results and evaluation reports; or in response to the switching command, it locks the region of interest and resets the segmentation weights, performs basic visual detection, and generates an evaluation report with maintenance prompts.
[0022] In this embodiment, Figure 1 The phrase "Constructing a multimodal time-series data stream" refers to the multimodal data alignment module. Figure 1 The “Generate Structure State Correlation Index” in the text refers to the correlation index generation module; Figure 1The phrase "Execute robust downgrade decision based on data confidence" in the text refers to the confidence decision downgrade module. Figure 1 The phrase "implementing visual linkage recognition and adaptive evaluation" refers to the visual linkage evaluation module. Example
[0023] The timing alignment and registration performed by the multi-modal data alignment module includes: Data acquisition is triggered using a precise time protocol to obtain the original vibration waveform and sound waveform with nanosecond-level timestamps; a short-time Fourier transform is performed on the original vibration waveform to extract the energy spectral density feature sequence, and the offset feature sequence of the first arrival time of the echo is extracted from the sound waveform. The energy spectral density feature sequence and the offset feature sequence are then resampled to a unified target frequency.
[0024] The timing alignment and registration performed by the multi-modal data alignment module also includes: A constraint window based on diagonal bandwidth is constructed. Within the constraint window, the Euclidean distance matrix between the energy spectral density feature sequence and the offset feature sequence is calculated. The dynamic time warping algorithm is used to find the warping path with the minimum cumulative distance. A one-to-one mapping relationship between time points is established based on the warping path.
[0025] In this embodiment, the dynamic time warping algorithm uses the Sakoe-Chiba band as the constraint window; the bandwidth parameter w of the constraint window is defined as the maximum allowable timing deviation window; the bandwidth parameter w is not arbitrarily set, but is stored in a JSON object in the system configuration file and loaded by the system initialization program before running; in this embodiment, considering that the propagation speed of sound waves in concrete medium is approximately 3000m / s to 4500m / s, and for prestressed concrete box girders with spans of 30m to 50m, the maximum physical path delay of sound wave reflection is usually no more than 500 milliseconds; at the same time, considering that the data sampling rate is 20Hz (i.e., one sampling point every 50 milliseconds), in order to cover the maximum physical delay and filter out irrelevant interference, the bandwidth parameter w is preferably set to 10 sampling points (corresponding to 500 milliseconds). The calculation logic for the bandwidth parameter w is as follows: when constructing the distance matrix, the algorithm only calculates the distance values of matrix elements (i, j) that satisfy the condition |i−j|≤w. For elements outside this range, their distance values are set to positive infinity, thereby forcing the regularized path to fall within the band-shaped region near the diagonal. Here, i and j represent the temporal indices of the two sets of feature sequences to be aligned (i.e., the energy spectral density feature sequence and the offset feature sequence) on their respective time axes, which are used to uniquely locate the local distance calculation unit between the feature points of the corresponding time steps in the two sequences. The correlation index generation module includes: Using the dual-channel long short-term memory network encoder in the cross-modal dynamic correlation feature fusion model, the modal vibration features and acoustic delay features in the multi-source heterogeneous time series dataset are encoded into vibration latent state vectors and acoustic latent state vectors, respectively; based on the interaction between the vibration latent state vectors and the acoustic latent state vectors at each time step, a collaborative gain weight is calculated to characterize the degree of temporal resonance between the two. The vibration latent state vector and the acoustic latent state vector are weighted and aggregated using the cooperative gain weight to generate the structural state correlation index, which represents the joint probability of bridge structural stiffness degradation and internal damage propagation within the current time window.
[0026] Furthermore, in the specific weighted aggregation calculation logic, the cross-modal dynamic correlation feature fusion model spatially aligns and concatenates the vibration latent state vector and the acoustic latent state vector through matrix operations, and assigns a cooperative gain weight calculated in real time by a normalized exponential function to achieve dynamic adjustment of the importance of different modal signals. For example, when the system detects a momentary heavy load impact on the bridge, if the vibration signal exhibits severe amplitude fluctuations and the acoustic signal simultaneously shows a reflection time delay shift, the normalized exponential function will output a higher cooperative gain weight (example value such as 0.88), making the fusion process highly focused on the coupling characteristics of these two; if noise interference in a certain mode weakens the correlation, the weight is automatically reduced (example value set to 0.12). Through this dynamic weighting based on energy interaction, random noise of a single mode is filtered out, structural variation features with resonance characteristics are retained and enhanced, and cross-modal fusion features containing deep coupling information are output. The acquisition of the joint probability of internal damage propagation is based on a deep integration of physical representation and statistical mapping: the system maps high-dimensional cross-modal fusion features to a one-dimensional scalar space through the fully connected layer, generating a feature value that comprehensively reflects the structural deterioration trend. Then, an S-shaped growth curve function (such as the Sigmoid function) is used to map this feature value to the interval [0,1]. In this mapping relationship, the output value is the joint probability: values close to 0 represent a stable structural period, while values close to 1 indicate that bridge stiffness degradation (manifested by vibration modal features) and internal crack / void propagation (manifested by acoustic delay features) have achieved a high degree of consistency in the spatiotemporal dimensions as damage criteria. This probability value directly quantifies the possibility of multiple defects coexisting and mutually inducing each other, providing quantitative guidance for the subsequent visual assessment module based on a multi-source feature evidence chain.
[0027] The correlation index generation module further includes: The cooperative interaction energy of the vibration hidden state vector and the sound wave hidden state vector at the current time step is calculated based on the hyperbolic tangent activation function; the cooperative interaction energy is normalized by the normalized exponential function to generate the cooperative gain weight with a value range between 0 and 1. The vibration latent state vector and the acoustic latent state vector are concatenated, and the concatenated vector is weighted and aggregated using the cooperative gain weight to obtain the cross-modal fusion feature. The cross-modal fusion feature is then dimensionality-reduced through a fully connected layer and mapped using an S-shaped growth curve function to finally generate the structural state correlation index.
[0028] In this embodiment, the generation of the cooperative gain weight adopts a gating mechanism based on the sigmoid growth curve function. The specific calculation logic is as follows: obtain the cooperative interaction energy scalar value et calculated at the current time step t; input the cooperative interaction energy scalar value (hereinafter denoted as et) into the standard sigmoid growth curve function and execute the following operation logic: Calculate the negative et power of the natural constant e, add 1 to the result, and finally take its reciprocal. This calculation process outputs a dimensionless real number in the range (0, 1), which is used as the cooperative gain weight at at time t. Furthermore, the cooperative gain weight at is a dimensionless scalar generated by normalization through an S-shaped growth curve function, with a theoretical range of (0, 1). When at approaches 0, i.e., a minimum trend, it means that at the current time t, the cooperative interaction energy et of the vibration latent state vector (hereinafter referred to as Hvib) and the acoustic latent state vector (hereinafter referred to as Haco) in the latent space is low, or even negative. This corresponds to two physical situations: one is that the system is in a completely healthy silent state with only ambient white noise; the other is that a single sensor has an outlier fault, including the acoustic probe falling off and causing false high-frequency signals, but the vibration sensor reading is normal. In this case, it is determined that there is no causal relationship between the two. In order to prevent single-source noise from falsely triggering the alarm, the information transmission at the current time is suppressed. When at approaches 1, i.e., when it is at a maximum value, it means that the vibration characteristics and the sound wave characteristics show a strong positive correlation and resonance. Physically, this corresponds to the moment when the internal damage and overall stiffness of the bridge structure decrease, which leads to the sudden change in sound wave time delay and the change in vibration mode occur simultaneously. At this time, it is determined that the two are causal evidence for each other, the characteristics at this moment are retained with high weight, and subsequent visual linkage is triggered. The core output cooperative gain weight at is mainly driven by the cooperative interaction energy et of the input parameter, which is determined by the vibration latent state vector Hvib and the acoustic latent state vector Haco. As et increases, at increases monotonically. et is calculated by the hyperbolic tangent function. Only when the directions of Hvib and Haco in the feature space tend to be consistent, that is, when damage coupling occurs, will the vector dot product produce a large positive value et. The bandwidth parameter w is nonlinearly constrained, and w determines the search range during dynamic time warping alignment. Setting w to 10 sampling points (500ms) is based on the physical propagation characteristics of sound waves in concrete media (3000-4500m / s). Too small a w will miss the real long-distance echo delay, leading to feature misalignment and thus reducing the calculated et and at. Too large a w will introduce irrelevant far-end noise, which will also interfere with the accuracy of et. Locking w within a physically reasonable range ensures that the feature sequences input to the fusion model are physically aligned on the time axis, guaranteeing the accurate calculation of the collaborative gain weights.
[0029] The correlation index generation module includes: Using the dual-channel long short-term memory network encoder in the cross-modal dynamic correlation feature fusion model, the modal vibration features and acoustic delay features in the multi-source heterogeneous time series dataset are encoded into vibration latent state vectors and acoustic latent state vectors, respectively; based on the interaction between the vibration latent state vectors and the acoustic latent state vectors at each time step, a collaborative gain weight is calculated to characterize the degree of temporal resonance between the two. The vibration latent state vector and the acoustic latent state vector are weighted and aggregated using the cooperative gain weight to generate the structural state correlation index, which represents the joint probability of bridge structural stiffness degradation and internal damage propagation within the current time window.
[0030] The correlation index generation module further includes: The cooperative interaction energy of the vibration hidden state vector and the sound wave hidden state vector at the current time step is calculated based on the hyperbolic tangent activation function; the cooperative interaction energy is normalized by the normalized exponential function to generate the cooperative gain weight with a value range between 0 and 1. The vibration latent state vector and the acoustic latent state vector are concatenated, and the concatenated vector is weighted and aggregated using the cooperative gain weight to obtain the cross-modal fusion feature. The cross-modal fusion feature is then dimensionality-reduced through a fully connected layer and mapped using an S-shaped growth curve function to finally generate the structural state correlation index.
[0031] Please see Figure 3 The robust degradation decision executed by the confidence decision degradation module includes: The information entropy of the modal vibration signal and the acoustic wave reflection time delay spectrum signal is calculated as the input confidence. This includes calculating the normalized information entropy of the modal vibration signal and the normalized information entropy of the acoustic wave reflection time delay spectrum signal for the signal time window at the current moment, and taking the reciprocal of the normalized information entropy of the two as the vibration input confidence and the acoustic wave input confidence, respectively.
[0032] The confidence decision degradation module compares the data integrity and reliability of the modal vibration signal and the acoustic reflection time delay spectrum signal based on a preset input confidence threshold. According to the comparison result, it determines whether the system is in a state of partial sensor failure or complete physical sensor failure, and triggers the corresponding soft degradation strategy or hard switching strategy accordingly. The input confidence threshold includes a vibration confidence benchmark threshold and an acoustic position confidence benchmark threshold. The classification decision includes executing the soft degradation strategy and the hard switching strategy.
[0033] In the confidence decision degradation module, when the input confidence of any signal is lower than the preset input confidence threshold but the other signal is normal, the function degradation maintenance in the case of partial sensor failure will execute a soft degradation strategy: in the cross-modal dynamic correlation feature fusion model, the attention weight corresponding to the low confidence signal will be forcibly set to zero. When both the vibration input confidence and the acoustic input confidence are below the input confidence threshold, a hard switching strategy is executed for the detection mode safety transfer in the event of full sensor failure. This includes cutting off the data path of the structural state correlation index and generating a switching command containing the visual dominant mode identifier bit and sending it to the visual linkage evaluation module.
[0034] Furthermore, in the confidence decision downgrade module, a normal signal means that the real-time calculated input confidence level is greater than or equal to the corresponding preset input confidence level threshold.
[0035] When the visual linkage evaluation module responds to the effective guidance information, it specifically performs the following linkage control: parses the numerical strength of the structural state correlation index, calculates the center coordinates and cropping size of the region of interest of the visual detection model on the input image through a preset spatial mapping function, and realizes optical focusing on the damaged area; establishes a positive correlation mapping relationship between the value of the structural state correlation index and the weight of the loss function in the image semantic segmentation algorithm, and dynamically increases the pixel-level classification weight of the crack category as the value of the structural state correlation index increases, so as to suppress the interference of background texture noise on the identification of fine cracks.
[0036] In this embodiment, a positive correlation mapping model between the structural state correlation index and the weights of the loss function in the image semantic segmentation algorithm is established to achieve accurate capture of minor defects in complex environments. Specifically, the visual linkage evaluation module uses a weighted cross-entropy loss function when performing semantic segmentation inference, and uses the real-time generated structural state correlation index as a dynamic adjustment factor to adjust the pixel-level classification weights of crack categories in real time through a preset mapping function. When the correlation index generation module calculates a high probability of damage based on vibration and sound wave signals, it automatically enhances the sensitivity of the positive correlation mapping model to crack features. Even in edge cases with low image contrast, extremely fine cracks, or severe texture interference (including bridge shadows, water seepage marks, and moss noise), the missed detection penalty coefficient of crack pixels in the loss function is dynamically increased, forcing the positive correlation mapping model to extract deep features from suspected pixels, thereby effectively separating hidden cracks from background noise. Upon receiving the switching command, the visual linkage evaluation module immediately interrupts data communication with the cross-modal dynamic correlation feature fusion model and forcibly locks the current input state of the visual detection model; it resets and locks the processing range of the visual detection model to the full field of view of the camera, and stops the dynamic coordinate adjustment based on the structural state correlation index to ensure that the maximum detection range is covered in the absence of prior guidance information. The classification weights of each category in the image semantic segmentation algorithm are restored to the preset initial equilibrium value, eliminating the weight bias for crack categories previously introduced by the high value of the structural state correlation index, and preventing overfitting false alarms when there is no internal data support. Based on the full field of view and the initial equalization value, the visual inspection model is run independently to identify cracks in the acquired images and generate a bridge health assessment report that includes a sensor data stream missing status identifier and the current visual inspection results.
[0037] Furthermore, the spatial mapping function is a homogeneous transformation matrix established based on the camera's intrinsic parameters (i.e., focal length and principal point) and extrinsic parameters (installation pose relative to the bridge coordinate system). It projects the physical space coordinates of the sensing module onto the two-dimensional image pixel plane, thereby calculating the center coordinates (u, v) of the region of interest (ROI) in real time based on the location of the signal anomaly source. The construction logic of the visual detection model is as follows: a semantic segmentation architecture based on a deep convolutional neural network is adopted. The encoder extracts multi-scale texture features of the bridge surface, and the decoder performs feature restoration by category in the pixel dimension, ultimately achieving accurate segmentation of defects including cracks and spalling.
[0038] Furthermore, the logic for dynamically increasing the weights is as follows: the structural state correlation index received in real time is used as an adjustment factor and applied to the loss function of the image semantic segmentation algorithm in real time; when the value of the structural state correlation index increases, the system automatically increases the penalty weight coefficient of the crack category in the loss function, making the visual model more sensitive to the pixel features of suspected cracks during backpropagation and inference; that is, since the internal physical index (structural state correlation index) has predicted the possibility of damage, the visual model will reduce its tolerance to false noise such as environmental shadows or moss, thereby extracting the edges of subtle cracks more accurately in low-contrast environments.
[0039] Furthermore, the visual inspection model refers to the specific mathematical entities and software components, including the architecture of the neural network, the stored weight parameter files, and the specific algorithm program that performs pixel-level classification; while the visual inspection mode refers to the working state and operation strategy of the system. In the linkage mode, the visual inspection model is guided by internal sensor data; while in the vision-dominated mode, the visual inspection model runs independently. Furthermore, the operational logic of restoring the initial equilibrium value and eliminating weight bias is to achieve neutral recovery of the algorithm when the physical sensing link fails: upon receiving a switching instruction, the system resets the previously amplified crack category weight coefficients in the loss function to the initial equilibrium state of 1:1; without an internal damage index as a confidence basis, continuing to maintain high weights will cause the model to overfit the background texture of any suspected lines, thus causing serious false alarms; by restoring the equilibrium value, the visual detection model reverts to a general detection state based on the image features themselves, performing blind patrols with the full field of view, ensuring that robust and unbiased detection results can still be provided when internal data references are lost.
[0040] In this embodiment, the input confidence level (hereinafter referred to as Cin) is calculated from the reciprocal of the normalized information entropy (hereinafter referred to as Hnorm), and its theoretical value range is [0, 1]. When Cin approaches 0, i.e., a minimum trend, it corresponds to Hnorm→1, i.e., maximum entropy. This means that the sensor signal presents a completely random and disordered state, without containing any definite structural information, characterizing sensor physical failure or extreme environmental interference, i.e., a floating level caused by a cable break or a strong electromagnetic pulse that drowns out the effective signal; it represents an unreliable state. If feature fusion is forcibly performed at this time, not only will effective features not be extracted, but serious noise pollution will be introduced, leading to distortion of the SSI index.
[0041] When Cin approaches 1, i.e., when it reaches a maximum value, it corresponds to Hnorm→0, i.e., low entropy. This means that the signal has a high degree of determinism and regularity, including clear vibration mode waveforms or sharp acoustic reflection peaks. It represents a high-confidence state, at which point the signal integrity is high, and the data source is determined to be sufficient to support subsequent cross-modal correlation calculations.
[0042] The input confidence level Cin comprises the vibration input confidence level Cvib and the acoustic input confidence level Caco; and the input confidence level Cin is directly controlled by the joint state of the input parameters vibration input confidence level (hereinafter referred to as Cvib) and acoustic input confidence level (hereinafter referred to as Caco). The normalized information entropy Hnorm is negatively correlated with the input confidence level Cin. According to Shannon's information theory, entropy is a measure of uncertainty. In structural health monitoring, effective signals (i.e., echoes generated by cracks) always exhibit energy concentration at a specific frequency or time point (low entropy), while faults or noise exhibit energy dispersion across the entire frequency band (high entropy). Therefore, using the reciprocal of entropy as the confidence level allows for the automatic identification and isolation of low-quality data, preventing the collapse of the entire assessment system due to the failure of a single sensor. The size of the region of interest (hereinafter referred to as SROI) and the input confidence Cin have a non-linear conditional dependency; when the system is in fusion mode, SROI is negatively correlated with SSI (focus); when the system is in switching mode, SROI is forcibly reset to the maximum value of the entire field of view; Furthermore, to verify the effectiveness of the confidence decision downgrade module and the visual linkage evaluation module, a full-system fault injection simulation platform was constructed. The physical object is a digital twin model of a prestressed concrete box girder. The signal generator generates standard vibration and acoustic sequences and can inject faults as needed, including drift faults, interruption faults, and noise faults. Specifically, the drift fault is superimposed with a linearly increasing DC component, the interruption fault is forced to set the signal amplitude to zero, and the noise fault is superimposed with high-power Gaussian white noise, i.e., SNR=-5dB. Table 1: Comparison of System Operation Modes and Performance Indicators under Different Sensor Confidence Levels Experimental scenario Hvib Haco Cvib Caco Mode Fco Bseg Sys-Perf Base-Perf Scene 1 0.35 0.42 0.85 0.78 Full-feature integration 15% 2.5 0.94 0.82 Scene 2 0.98 0.41 0.12 0.79 Soft downgrade 35% 1.8 0.88 0.45 Scene 3 0.36 0.95 0.84 0.15 Soft downgrade 30% 1.6 0.86 0.52 Scene 4 0.99 0.97 0.11 0.13 Hard switch 100% 1.0 0.78 0.15 Scene 5 0.38 0.45 0.82 0.75 Full-feature integration 18% 2.2 0.92 0.80 Scene Six 0.92 0.40 0.18 0.80 Soft downgrade 32% 1.9 0.87 0.48 In Table 1 above, Hvib and Haco are used to quantify the "disorder" or "noise content" of sensor signals. The numerical range is [0, 1]; the closer the value is to 0, the clearer and more regular the signal characteristics (high quality); the closer the value is to 1, the closer the signal is to random noise, and the less effective information is available (low quality / fault).
[0043] Cvib and Caco are derived parameters, representing input confidence levels. Cvib represents the confidence level of vibration input, and Caco represents the confidence level of sound wave input. They are calculated based on the inverse relationship of information entropy. Specifically, Hvib and Haco are extracted from the signal to characterize the degree of signal disorder, and then confidence indices Cvib and Caco, which are inversely proportional, are established through mathematical transformations. The calculation logic of both is compatible; the calculation logic of Hvib and Haco is responsible for measuring disorder, while the calculation logic of Cvib and Caco is responsible for measuring usability. They are mutually complementary and together constitute... The criteria for the confidence decision degradation module are as follows: The input confidence level directly determines the system's level of trust in the signal; the higher the value, the closer it is to 1, the more reliable the sensor data is; if the value is lower than the preset input confidence threshold, i.e., 0.4, the sensor is determined to be faulty; Mode represents the operating mode, i.e., the robust control strategy currently activated by the system; Based on the joint state of Cvib and Caco, the system automatically selects the decision path, including "full-function fusion" (best performance), "soft degradation" (single-path fault tolerance), and "hard switching" (safety fallback). Fcov represents the visual coverage factor, which is the percentage of coverage of the visual detection model; it is used to characterize the focus of the visual system. 15% represents high-precision local magnification focus, which relies on precise guidance; 100% represents undifferentiated full-field wide-angle scanning, which is a safe mode when guidance is lacking. Bseg is the classification weight multiple of the semantic segmentation algorithm for crack categories, representing the linked output, i.e., visual weight bias; a value of 1.0 represents balanced detection; a value greater than 1.0 indicates that the system is extremely sensitive to minute crack features under the guidance of prior information. Sys−Perf and Base−Perf are detection accuracy. Sys−Perf is the F1-Score of this invention, which is the harmonic mean of precision and recall. Base−Perf is the F1-Score of the control group (benchmark system); they are used to measure the overall reliability of the detection results. The visual coverage factor Fcov is a derived parameter used to quantify the search range of the vision system; the lower the value, the more accurate the focus, and a value of 100% indicates full field of view scanning; the visual weight bias Bseg is a derived parameter used to quantify the degree of attention of the semantic segmentation algorithm to the crack category, and Bseg=1.0 indicates unbiased equal detection.
[0044] The six scenarios in Table 1 simulate different sensor health states to verify the adaptability and robustness of the invention under various extreme conditions: In scenarios one and five, normal monitoring and full-function fusion are performed. At this time, the entropy values of vibration and acoustic signals are low (H≈0.35−0.45), and the confidence levels are high (C>0.75). This indicates that the system is in an ideal working environment and all sensors are operating well. At this time, the system enters the full-function fusion mode. Due to the high reliability of the data, the system directs the vision module to perform high-precision focusing (Fcov≈15%−18%) and significantly improves the crack recognition weight (Bseg>2.0). The performance of this invention (0.94 / 0.92) is significantly better than that of the control group (0.82 / 0.80), which proves that under normal conditions, multimodal fusion can effectively improve the detection limit.
[0045] In scenarios two and six, the vibration sensor malfunctions, i.e., single-point failures. The vibration signal entropy spikes to nearly 1.0 (Hvib≈0.98 and 0.92), and the confidence level plummets below the threshold (Cvib≈0.12 and 0.18), while the acoustic signal remains normal. This simulates a scenario where the vibration sensor detaches or the cable has poor contact. The system identifies the single-path failure and triggers a soft degradation strategy. At this point, interference from the vibration data is automatically eliminated, and guidance is based solely on the acoustic data. The visual focus is moderately relaxed (Fcov≈32%−35%) to tolerate a slight decrease in guidance accuracy. This invention still maintains high performance (0.88 / 0.87). In contrast, the control group, lacking a degradation mechanism, incorrectly integrates noise data, resulting in a catastrophic performance decline (0.45 / 0.48). Scenario 3 is a single-point failure of the acoustic probe, which is the opposite of Scenario 2. In this case, the entropy of the acoustic signal is high (Haco=0.95) and the confidence level is low (Caco=0.15), while the vibration signal is normal, simulating the situation of acoustic probe damage. The system also triggers the soft degradation strategy, but instead removes the acoustic data and relies only on the vibration data. This proves that the degradation mechanism of the present invention is asymmetric and flexible. No matter which sensor fails, the system can maintain a performance level of more than 85% (0.86), while the control group fails due to noise pollution (0.52).
[0046] Scenario 4 represents a complete sensor failure, an extreme disaster. The entropy values of both vibration and acoustic signals approach 1.0, with low confidence levels (≈0.1). This simulates extreme situations where all peripheral sensors become unusable due to strong electromagnetic storms, power outages of the acquisition card, or communication bus disconnections. The system determines that external guidance information is completely unreliable, triggering a hard switching strategy. The system disconnects the fusion module, forcing the vision system into "independent mode": full field of view (Fcov=100%), and weights are restored to equilibrium (Bseg=1.0). This invention successfully achieves a "safety fallback," maintaining a passable performance of 0.78 based solely on pure vision, ensuring basic detection capabilities. In contrast, the control group system completely collapsed (0.15), losing all usability. In Scenario 4, the low Base-Perf (0.15) of the control group is due to the fact that traditional hard fusion algorithms, when all sensors fail, misjudge the input pure noise signal as high-intensity damage features, thus calculating incorrect ROI coordinates (random jumps) and applying high incorrect segmentation weights. This causes the visual model to focus on the background area and generate a large number of false crack detections, resulting in an accuracy rate close to zero.
[0047] Furthermore, the vibration input confidence Cvib and acoustic input confidence Caco directly reflect the front-end health of the system's perception, and monitoring them is the most effective means of determining the system's back-end strategy (i.e., fusion or independent vision).
[0048] Through large-scale experimental statistics, a mapping relationship between signal confidence, signal-to-noise ratio (SNR), and information entropy was established. The input confidence threshold was locked at a physical critical point of -3dB SNR (i.e., Zth=0.4), which served as an objective boundary for judging signal validity. This allowed for the timely identification and removal of low SNR noise data that had lost fusion value before the feature extraction bit error rate increased exponentially. When Cin < 0.4 (corresponding to normalized information entropy Hnorm > 2.5), the signal SNR was below -3dB. At this point, the feature extraction bit error rate increased exponentially, and the data lost its fusion value. Therefore, 0.4 is the objective physical boundary for distinguishing usable signals from noise garbage.
[0049] It should be noted that all calculation formulas in this application employ regression analysis, including but not limited to machine learning algorithms, to deeply analyze the collected parameters and identify their natural trends and interrelationships. Specialized software, such as Python's Scikit-learn library or the R language, is used to automatically generate mathematical models that match the data. Then, cross-validation and other methods are used to objectively evaluate the model performance, and continuous feedback and optimization are combined to ensure that the created formulas truly reflect the inherent laws of the data, thereby guaranteeing their effectiveness and accuracy. In all calculation formulas in this application, the parameters in each formula undergo dimensionless processing within a consistent range to ensure that different physical quantities are compared on the same scale; dimensionless processing techniques include, but are not limited to, min-max-normalization and Z-score standardization. The technical solution of this invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random-access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments of this invention.
[0050] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0051] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A machine vision-based intelligent bridge condition detection and evaluation system, characterized in that, The specific steps include: The multi-mode data alignment module is used to synchronously acquire the modal vibration signals and acoustic wave reflection time delay spectrum signals of the bridge structure. Through a high-precision timestamp synchronization mechanism and a dynamic time warping algorithm, the modal vibration signals and acoustic wave reflection time delay spectrum signals are time-series aligned and registered within a unified time window to generate a multi-source heterogeneous time series dataset. The correlation index generation module inputs the multi-source heterogeneous time-series dataset into a preset cross-modal dynamic correlation feature fusion model; the cross-modal dynamic correlation feature fusion model is constructed based on a multi-task learning framework, and uses a recurrent neural network to extract vibration mode features and acoustic delay features respectively, and calculates the structural state correlation index; The confidence decision degradation module is used to monitor the input confidence of modal vibration signals and acoustic wave reflection time delay spectrum signals in real time, and trigger a robust degradation mechanism when either input confidence is lower than a preset input confidence threshold. This includes automatically adjusting the feature fusion weights in the cross-modal dynamic correlation feature fusion model to reduce the impact of low-confidence signals, or blocking the output of the structural state correlation index and generating a switching command to control the system to enter a single visual detection mode; otherwise, the structural state correlation index is output as effective guidance information. The visual linkage evaluation module, in response to the effective guidance information, dynamically adjusts the coordinates of the region of interest and the semantic segmentation weights of the visual detection model based on the structural state correlation index, and generates refined recognition results and evaluation reports; or in response to the switching command, it locks the region of interest and resets the segmentation weights, performs basic visual detection, and generates an evaluation report with maintenance prompts.
2. The intelligent bridge condition detection and evaluation system based on machine vision according to claim 1, characterized in that: The timing alignment and registration performed by the multi-modal data alignment module includes: Data acquisition is triggered using a precise time protocol to obtain the original vibration waveform and sound waveform with nanosecond-level timestamps; a short-time Fourier transform is performed on the original vibration waveform to extract the energy spectral density feature sequence, and the offset feature sequence of the first arrival time of the echo is extracted from the sound waveform. The energy spectral density feature sequence and the offset feature sequence are then resampled to a unified target frequency.
3. The intelligent bridge condition detection and evaluation system based on machine vision according to claim 2, characterized in that: The timing alignment and registration performed by the multi-modal data alignment module also includes: A constraint window based on diagonal bandwidth is constructed. Within the constraint window, the Euclidean distance matrix between the energy spectral density feature sequence and the offset feature sequence is calculated. The dynamic time warping algorithm is used to find the warping path with the minimum cumulative distance. A one-to-one mapping relationship between time points is established based on the warping path.
4. The intelligent bridge condition detection and evaluation system based on machine vision according to claim 3, characterized in that: The correlation index generation module includes: Using the dual-channel long short-term memory network encoder in the cross-modal dynamic correlation feature fusion model, the modal vibration features and acoustic delay features in the multi-source heterogeneous time series dataset are encoded into vibration latent state vectors and acoustic latent state vectors, respectively; based on the interaction between the vibration latent state vectors and the acoustic latent state vectors at each time step, a collaborative gain weight is calculated to characterize the degree of temporal resonance between the two. The vibration latent state vector and the acoustic latent state vector are weighted and aggregated using the cooperative gain weight to generate the structural state correlation index, which represents the joint probability of bridge structural stiffness degradation and internal damage propagation within the current time window.
5. The intelligent bridge condition detection and evaluation system based on machine vision according to claim 4, characterized in that: The correlation index generation module further includes: The cooperative interaction energy of the vibration hidden state vector and the sound wave hidden state vector at the current time step is calculated based on the hyperbolic tangent activation function; the cooperative interaction energy is normalized by the normalized exponential function to generate the cooperative gain weight with a value range between 0 and 1. The vibration latent state vector and the acoustic latent state vector are concatenated, and the concatenated vector is weighted and aggregated using the cooperative gain weight to obtain the cross-modal fusion feature. The cross-modal fusion feature is then dimensionality-reduced through a fully connected layer and mapped using an S-shaped growth curve function to finally generate the structural state correlation index.
6. The intelligent bridge condition detection and evaluation system based on machine vision according to claim 5, characterized in that: The robust degradation decision executed by the confidence decision degradation module includes: The information entropy of the modal vibration signal and the acoustic wave reflection time delay spectrum signal is calculated as the input confidence. This includes calculating the normalized information entropy of the modal vibration signal and the normalized information entropy of the acoustic wave reflection time delay spectrum signal for the signal time window at the current moment, and taking the reciprocal of the normalized information entropy of the two as the vibration input confidence and the acoustic wave input confidence, respectively.
7. The intelligent bridge condition detection and evaluation system based on machine vision according to claim 6, characterized in that: The confidence decision degradation module compares the data integrity and reliability of the modal vibration signal and the acoustic reflection time delay spectrum signal based on a preset input confidence threshold. According to the comparison result, it determines whether the system is in a state of partial sensor failure or complete physical sensor failure, and triggers the corresponding soft degradation strategy or hard switching strategy accordingly. The input confidence threshold includes a vibration confidence benchmark threshold and an acoustic position confidence benchmark threshold. The classification decision includes executing the soft degradation strategy and the hard switching strategy.
8. The intelligent bridge condition detection and evaluation system based on machine vision according to claim 7, characterized in that: In the confidence decision degradation module, when the input confidence of any signal is lower than the preset input confidence threshold but the other signal is normal, the function degradation maintenance in the case of partial sensor failure will execute a soft degradation strategy: in the cross-modal dynamic correlation feature fusion model, the attention weight corresponding to the low confidence signal will be forcibly set to zero. When both the vibration input confidence and the acoustic input confidence are below the input confidence threshold, a hard switching strategy is executed for the detection mode safety transfer in the event of full sensor failure. This includes cutting off the data path of the structural state correlation index and generating a switching command containing the visual dominant mode identifier bit and sending it to the visual linkage evaluation module.
9. The intelligent bridge condition detection and evaluation system based on machine vision according to claim 8, characterized in that: When the visual linkage evaluation module responds to the effective guidance information, it specifically performs the following linkage control: parses the numerical strength of the structural state correlation index, calculates the center coordinates and cropping size of the region of interest of the visual detection model on the input image through a preset spatial mapping function, and realizes optical focusing on the damaged area; establishes a positive correlation mapping relationship between the value of the structural state correlation index and the weight of the loss function in the image semantic segmentation algorithm, and dynamically increases the pixel-level classification weight of the crack category as the value of the structural state correlation index increases, so as to suppress the interference of background texture noise on the identification of fine cracks.
10. The intelligent bridge condition detection and evaluation system based on machine vision according to claim 9, characterized in that: Upon receiving the switching command, the visual linkage evaluation module immediately interrupts data communication with the cross-modal dynamic correlation feature fusion model and forcibly locks the current input state of the visual detection model; it resets and locks the processing range of the visual detection model to the full field of view of the camera, and stops the dynamic coordinate adjustment based on the structural state correlation index to ensure that the maximum detection range is covered in the absence of prior guidance information. The classification weights of each category in the image semantic segmentation algorithm are restored to the preset initial equilibrium value, eliminating the weight bias for crack categories previously introduced by the high value of the structural state correlation index, and preventing overfitting false alarms when there is no internal data support. Based on the full field of view and the initial equalization value, the visual inspection model is run independently to identify cracks in the acquired images and generate a bridge health assessment report that includes a sensor data stream missing status identifier and the current visual inspection results.
Citation Information
Patent Citations
Bridge health state detection and early warning method and system based on image recognition
CN119888496A