A method and device for monitoring bridge crack width
By obtaining video images of bridge surfaces in real time and using object detection models and time series prediction models, real-time monitoring of bridge cracks and prediction of width change trends is achieved, solving the problem of insufficient real-time and accuracy in traditional methods, and providing an efficient bridge health monitoring solution.
Patent Information
- Application Number
- CN202510143671.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-10
AI Technical Summary
Traditional bridge crack monitoring methods are difficult to achieve real-time and high-precision crack data acquisition, and the monitoring efficiency is inefficient.
By acquiring continuous video images of the bridge surface, using a pre-trained object detection model for feature extraction, combining optical flow calculation algorithms and time series prediction models, real-time monitoring of cracks and width change trend prediction are achieved.
It realizes high-precision crack detection and analysis, provides key dynamic information for bridge health monitoring, can early warning of potential structural risks, and provides a scientific basis for long-term maintenance and repair of bridges.
Smart Images

Figure CN119600029B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method and device for monitoring the width of bridge cracks. Background Art
[0002] With the aging of bridge facilities and the influence of the external environment, cracks, as one of the main forms of damage to bridge structures, have become a key factor affecting bridge safety and life. Traditional bridge crack monitoring mostly relies on manual detection or static monitoring equipment, which makes it difficult to obtain high-precision crack data in real time and has low monitoring efficiency.
[0003] Therefore, how to use modern artificial intelligence technology to achieve real-time monitoring and health assessment of bridge cracks has become a research hotspot and technical challenge. Summary of the invention
[0004] The present application provides a bridge crack width monitoring method and device to achieve efficient and accurate crack width monitoring.
[0005] This application provides the following solutions:
[0006] According to the first method, a method for monitoring the width of a bridge crack is provided, the method comprising: obtaining a continuous video image of a bridge surface, the continuous video image comprising a plurality of frame images; loading a pre-trained target detection model, using the target detection model to perform feature extraction on each frame image in the continuous video image, obtaining image features of the frame image, identifying the location of the crack using the image features through a multi-layer convolutional neural network, and generating a segmentation mask map of the crack; wherein the target detection model is trained using a first data set, the first data set comprising a first frame image sample and its corresponding crack recognition result sample, when pre-training the target detection model, extracting optical flow features between adjacent frame images in the first frame image sample through an optical flow calculation algorithm, and merging the optical flow features with the image features of the frame image sample; extracting the contour pixel points of the crack based on the crack segmentation mask map, calculating the pixel normal direction width of the crack in combination with the image gradient information, and obtaining the width change information and width change rate of the crack; the width change information comprises: average width change value, local width degree change value and / or relative width change value; according to the width change information and the width transformation rate, the time series prediction model is used to predict the trend of the crack width change to obtain the prediction result of the crack width change; wherein, the time series prediction model uses the second data set for model training, the second data set includes the second frame image sample and its corresponding prediction result sample, in the training process of the time series prediction model, a basic model and a complexity capture module are designed, the basic model generates a first prediction result based on the input frame image sample; the complexity capture module performs complex area recognition based on the input frame image sample, and generates a second prediction result for the complex area; the first prediction result and the second prediction result are weightedly fused to obtain a third prediction result; the basic model and the complexity capture module are iterated with the goal of minimizing the difference between the third prediction result and the prediction result sample; wherein, in each round of parameter iteration, the parameters of the basic model are updated using a fixed learning rate, and the parameters of the complexity capture module are updated using a dynamic learning rate.
[0007] According to an achievable method in an embodiment of the present application, the method also includes: obtaining environmental parameter data of the bridge site, the environmental parameters including temperature, humidity, wind speed and light intensity; using the target detection model to extract features from each frame image in the continuous video image includes: inputting the environmental parameters and the video image into the target detection model at the same time; encoding the features of the environmental parameters through a multi-layer perceptron, and deeply fusing the encoded environmental features with the video image features; the first data set also includes environmental parameter samples; when pre-training the target detection model, the crack data in the first data set is grouped according to the environmental parameter samples to establish multiple sub-data sets; training multiple sub-models, each sub-model performs weight update based on the corresponding sub-data set; sharing the weights of the feature extraction layer among the multiple sub-models, and integrating the multiple sub-models to generate the target detection model.
[0008] According to an achievable method in an embodiment of the present application, the fusing of the optical flow features with the image features of the frame image samples includes: calculating the pixel motion vectors between consecutive frames through an optical flow algorithm to extract inter-frame motion characteristics; performing multi-layer convolution processing on the optical flow features to generate an optical flow feature map; utilizing a channel attention mechanism to fuse the optical flow feature map with the image features, and retaining the key features of the original image through a residual connection; and normalizing the fused feature map and inputting it into the target detection model.
[0009] According to an achievable method in an embodiment of the present application, after extracting the contour pixel points of the crack based on the crack segmentation mask image, it also includes: using a local curvature algorithm to calculate the curvature value of each point in the contour pixel points; identifying the corner points in the contour pixel points based on the curvature values of each point; and modifying the pixel points on both sides of the tangent direction of the corner point into contour pixel points through an interpolation algorithm.
[0010] According to an achievable method in an embodiment of the present application, the design of the basic model and the complexity capture module includes: the basic model adopts a single-layer long short-term memory network structure to construct a model; the complexity capture module adopts a double-layer long short-term memory network structure to locally model the drastic changes in crack width in high-complexity areas; during the training process of the time series prediction model, a complexity evaluation function is used to dynamically adjust the learning rate of the complexity capture module according to the crack width change rate; the output results of the basic model and the complexity capture module are combined, and the parameters of the time series prediction model are optimized through a weighted attention mechanism.
[0011] According to an achievable method in an embodiment of the present application, the method further includes: after acquiring continuous video images of the bridge surface, preprocessing the video images, and the preprocessing includes one or more of the following: grayscale processing of the frame images; using a combination of Gaussian filtering and edge-preserving filtering algorithms to remove noise in the frame images; using a contrast enhancement algorithm to enhance the contrast of the frame images; and normalizing the pixel values of the frame images.
[0012] According to an achievable method in an embodiment of the present application, the method further includes: evaluating the health status of the bridge using the predicted results of the crack width change trend, and sending the width change information, the width change rate, the predicted results of the crack width change and / or the bridge health status report to a monitoring center through an Internet of Things platform, wherein the monitoring center updates the health status assessment results in real time and displays the results visually; wherein the health status assessment includes: dividing the risk levels of bridge cracks based on the predicted results of the crack width change, and generating bridge maintenance recommendations based on different risk levels, including crack repair, local reinforcement or comprehensive inspection.
[0013] According to the second aspect, a bridge crack width monitoring device is provided, the device comprising: an image acquisition unit, configured to acquire continuous video images of a bridge surface, the continuous video images comprising a plurality of frame images; a target detection unit, configured to load a pre-trained target detection model, and use the target detection model to perform feature extraction on each frame image in the continuous video images, identify the location of the cracks through a multi-layer convolutional neural network, and generate a segmentation mask map of the cracks; wherein, when the target detection model is pre-trained, the optical flow features between adjacent frames in a first data set are extracted through an optical flow calculation algorithm, and the optical flow features are deeply fused with the image features; a width change analysis unit, configured to extract the contour pixel points of the crack based on the crack segmentation mask map, and calculate the pixel normal direction width of the crack in combination with the image gradient information. , obtain the width change information and width change rate of the crack; the width change information includes: an average width change value, a local width change value and / or a relative width change value; a width change prediction unit is configured to use a time series prediction model to perform trend prediction on the crack width change according to the width change information and the width change rate, and obtain a prediction result of the crack width change; wherein, the time series prediction model uses a second data set for model training, and during the training process of the time series prediction model, a basic model and a complexity capture module are designed, the basic model uses a fixed learning rate, and captures local subtle changes in the crack width based on an adaptive attention mechanism; the complexity capture model uses a dynamic learning rate in the complexity area; the results of the basic model and the complexity capture module are combined, and a weighted fusion strategy is used to optimize the weights of the main model.
[0014] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the steps of any one of the methods described in the first aspect are implemented.
[0015] According to a fourth aspect, an electronic device is provided, comprising: one or more processors; and a memory associated with the one or more processors, the memory being used to store program instructions, and the program instructions, when read and executed by the one or more processors, execute the steps of any one of the methods described in the first aspect.
[0016] According to the specific embodiments provided in this application, this application discloses the following technical effects:
[0017] (1) This application effectively extracts the pixel-level segmentation mask of the cracks by acquiring continuous video images of the bridge surface in real time and processing them using a pre-trained target detection model, accurately identifying the location and shape of the cracks, thereby achieving high-precision crack detection and analysis. The accuracy and stability of crack location detection are further improved by using deeply fused optical flow features and image features. By calculating the pixel normal width of the crack, the width change information and rate of the crack are obtained, providing key dynamic information for bridge health monitoring. This application also combines a time series prediction model to predict the changing trend of the crack width, which can provide early warning of potential structural risks and provide a scientific basis for the long-term maintenance and repair of bridges.
[0018] (2) This application introduces environmental parameter data. By inputting environmental parameters into the model together with video images, and using a multi-layer perceptron to encode environmental features and deeply integrate them with video image features, the interference caused by environmental factors on crack detection can be effectively reduced. Grouping crack data and training sub-models based on environmental parameter samples in the first data set helps the model adapt to crack changes under different environmental conditions, enhances the adaptability and universality of the detection model, and enables the method of this application to maintain good detection results in a changing actual environment.
[0019] (3) This application deeply fuses optical flow features with image features to effectively extract motion information and static features from continuous video frames, thereby improving the accuracy of crack detection. Feature fusion through the channel attention mechanism and the use of residual connections to retain the key features of the original image help retain important visual information while reducing feature loss caused by fusion. The normalized feature map is input into the target detection model, which further improves the model's ability to recognize cracks and its stability. This feature fusion method can sensitively capture subtle changes in cracks and ensure the accuracy and robustness of real-time monitoring.
[0020] (4) This application uses a local curvature algorithm to calculate the curvature value of contour points, which can effectively identify corner points and process them. It can accurately locate corner points in the crack contour, thereby avoiding the problem of inaccurate width calculation caused by corner point discontinuity or error. The interpolation algorithm is used to smooth the pixels on both sides of the tangent direction of the corner point, further optimizing the continuity of the contour, making the crack width calculation more stable and accurate in complex contour areas.
[0021] (5) The design of the basic model and complexity capture module used in the time series prediction model of this application can enhance the ability to predict the trend of crack width changes. The basic model uses a single-layer LSTM network to capture the overall trend, ensuring that the long-term trend of crack width changes can be captured. The complexity capture module uses a double-layer LSTM structure to achieve local modeling of areas with drastic changes, enhancing the model's sensitivity to rapid changes. The strategy of dynamically adjusting the learning rate using a complexity evaluation function enables the model to show better adaptability when dealing with areas of different complexity, thereby optimizing the accuracy of crack width prediction. The combined use of the weighted attention mechanism helps to combine the advantages of both and obtain more accurate prediction results.
[0022] (6) This application preprocesses continuous video images, including grayscale processing, Gaussian filtering, edge-preserving filtering, contrast enhancement, and pixel value normalization, which can significantly improve the quality of subsequent image analysis. The preprocessing process can remove noise, enhance the contrast and details of the image, reduce the errors caused by environmental influences, and enable the target detection model to extract crack features more accurately. Such preprocessing steps improve the reliability and robustness of the crack detection algorithm in various environments and conditions, and provide a basis for achieving real-time and accurate bridge monitoring.
[0023] (7) This application uses the prediction results of crack width change trends for bridge health status assessment, which can realize intelligent bridge monitoring and maintenance decision support based on real-time data. Using the Internet of Things platform to send health status reports and update assessment results in real time helps to achieve remote monitoring and dynamic display of bridge health status. By dividing the risk level of crack development and generating maintenance recommendations, different response measures can be taken according to different risk levels, thereby improving the safety and service life of the bridge. This method provides a data-driven bridge maintenance management system that improves risk warning capabilities and maintenance efficiency.
[0024] Of course, any product implementing the present application does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0026] Figure 1 A flow chart of a bridge crack width monitoring method provided in an embodiment of the present application;
[0027] Figure 2 is a real image of the image frame in the embodiment of the present application;
[0028] FIG3 (a) is a diagram showing the variation of the average width of cracks in an embodiment of the present application;
[0029] FIG3( b ) is a diagram showing the local width variation of a crack in an embodiment of the present application;
[0030] FIG3 (c) is a diagram showing the relative width variation of cracks in an embodiment of the present application;
[0031] Figure 4 A schematic block diagram of a bridge crack width monitoring device provided in an embodiment of the present application;
[0032] Figure 5 A schematic block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0033] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.
[0034] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.
[0035] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0036] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.
[0037] There are already some crack monitoring technologies. For example, the static image analysis method uses a fixed camera to capture periodic static images of the bridge surface for analysis. Although this method can capture visible structural changes, it cannot reflect dynamic changes in real time. The sensor monitoring method uses various sensors (such as strain gauges, accelerometers) to monitor changes in the mechanical parameters of the bridge. Although this method can deeply analyze the health status of the structure, its direct observation and measurement capabilities of cracks are limited, and the installation and maintenance costs are high. The image processing algorithm uses image processing technology based on traditional computer vision to detect cracks. This method has high requirements on image quality and algorithm robustness, and is often affected by noise and ambient light.
[0038] In view of this, the present application provides a new idea. Figure 1 A flow chart of the bridge crack width monitoring method provided in the embodiment of the present application is as follows: Figure 1 As shown in , the method may include the following steps:
[0039] Step 101: Acquire continuous video images of the bridge surface, where the continuous video images include multiple frame images.
[0040] Step 102: Load the pre-trained target detection model, use the target detection model to extract features from each frame image in the continuous video image, obtain image features of the frame image, identify the location of the crack using the image features through a multi-layer convolutional neural network, and generate a segmentation mask map of the crack; wherein the target detection model uses the first data set for model training, the first data set includes a first frame image sample and its corresponding crack identification result sample, and when pre-training the target detection model, extract the optical flow features between adjacent frame images in the first frame image sample through an optical flow calculation algorithm, and fuse the optical flow features with the image features of the frame image sample.
[0041] Step 103: extract the contour pixel points of the crack based on the crack segmentation mask image, calculate the pixel normal width of the crack in combination with the image gradient information, and obtain the width change information and width change rate of the crack; the width change information includes: average width change value, local width change value and / or relative width change value.
[0042] Step 104: According to the width change information and the width transformation rate, a time series prediction model is used to perform trend prediction on the crack width change to obtain a prediction result of the crack width change; wherein, the time series prediction model is trained using a second data set, and the second data set includes a second frame image sample and its corresponding prediction result sample. During the training process of the time series prediction model, a basic model and a complexity capture module are designed, and the basic model generates a first prediction result based on the input frame image sample; the complexity capture module identifies a complex area based on the input frame image sample, and generates a second prediction result for the complex area; the first prediction result and the second prediction result are weightedly fused to obtain a third prediction result; the basic model and the complexity capture module are iterated with the goal of minimizing the difference between the third prediction result and the prediction result sample; wherein, in each round of parameter iteration, the parameters of the basic model are updated using a fixed learning rate, and the parameters of the complexity capture module are updated using a dynamic learning rate.
[0043] It can be seen from the above process that this application effectively extracts the pixel-level segmentation mask map of the cracks by acquiring continuous video images of the bridge surface in real time and processing them using a pre-trained target detection model, accurately identifying the position and shape of the cracks, thereby achieving high-precision crack detection and analysis. The accuracy and stability of crack position detection are further improved by using deeply fused optical flow features and image features. By calculating the pixel normal width of the crack, the width change information and rate of the crack are obtained, providing key dynamic information for bridge health monitoring. This application also combines a time series prediction model to predict the changing trend of the crack width, which can provide early warning of potential structural risks and provide a scientific basis for the long-term maintenance and repair of bridges.
[0044] The following is a detailed description of each step in the above process and the effects that can be further produced in conjunction with the embodiments. It should be noted that the "first" and "second" and other limitations involved in the present disclosure do not have limitations in terms of size, order, and quantity, and are only used to distinguish them in name, for example, "first data set" and "second data set" are used to distinguish two data sets.
[0045] First, the above step 101, namely "obtaining continuous video images of the bridge surface, the continuous video images including a plurality of frame images" is described in detail in conjunction with the embodiment.
[0046] This application continuously captures image data of the bridge surface through high-resolution cameras or video monitoring equipment installed on the bridge surface or its periphery. The acquired video images are composed of multiple continuous frame images, which are captured and recorded at a high rate to ensure the real-time and continuity of monitoring. By analyzing the continuous video images, the changing characteristics of the bridge surface at different time points can be captured, especially in terms of the evolution of crack location and morphology. Each frame image in the continuous video stream contains detailed information about the current bridge status, Figure 2 It is a real image of the image frame in the embodiment of the present application, which provides a rich data source for subsequent crack detection and analysis.
[0047] In the specific implementation process, the acquisition of video images can be achieved through a variety of devices and technologies, such as fixed cameras, drones or other mobile devices, which can shoot at a high frame rate (such as 30 frames per second or higher) to ensure that subtle changes in the bridge surface can be captured.
[0048] The acquired video image serves as the basic input for subsequent target detection. In order to improve the accuracy of crack monitoring, the present application can preprocess the video image. The preprocessing includes one or more of the following: graying the frame image; removing noise in the frame image by combining Gaussian filtering and edge-preserving filtering algorithms; improving the contrast of the frame image by a contrast enhancement algorithm; and normalizing the pixel values of the frame image.
[0049] Grayscale processing is the process of converting a color image into a grayscale image. This process generates a single grayscale value by weighted averaging the red, green, and blue (RGB) color channel values of each pixel. Grayscale processing helps reduce the computational burden and makes subsequent image analysis and feature extraction more efficient. Grayscale processing helps to increase the processing speed of the image and provide clearer image input for subsequent feature extraction and analysis. In particular, when analyzing crack edges and details, removing color information helps reduce noise interference and improve the stability of the algorithm.
[0050] Gaussian filtering is a commonly used smoothing filtering technique that removes random noise by convolving an image with a Gaussian function. Its implementation includes calculating the weighted average of each pixel and its surrounding neighborhood pixels to achieve a smoothing effect. This process helps to eliminate image noise caused by environmental factors (such as wind, rain, etc.), making the key features in the image more obvious. In the present invention, in order to maintain the clarity of image edge features while denoising, an edge-preserving filtering algorithm is also combined. Edge-preserving filtering is a technique that aims to smooth an image while retaining edges and details. Common implementation methods include bilateral filtering and guided filtering. By combining Gaussian filtering with edge-preserving filtering, noise in the image can be effectively removed while ensuring that subtle features such as cracks are not blurred or lost, thereby improving the accuracy of crack detection.
[0051] The contrast enhancement algorithm aims to adjust the brightness and contrast of the image to make the features in the image more prominent and facilitate subsequent analysis. In the present invention, methods such as histogram equalization or adaptive histogram equalization can be used. Histogram equalization redistributes the grayscale of the image to make the brightness distribution of the image more uniform and enhance the contrast of the image. Adaptive histogram equalization (such as the CLAHE algorithm) can also be used to perform histogram equalization in local areas to address the problem of uneven brightness in different areas.
[0052] Normalization is to map the pixel values of an image to a specific range, usually 0 to 1 or 0 to 255, in order to unify the scale of the data and reduce the impact of brightness differences between different frames on the analysis results. In this application, normalization can be accomplished by subtracting the minimum value of the image from the pixel value and dividing it by the range value (the difference between the maximum and minimum values).
[0053] When acquiring continuous video images of the bridge surface, the present application can also synchronously collect environmental parameter data of the bridge site, wherein the environmental parameters include temperature, humidity, wind speed and light intensity. The environmental information is used to assist the subsequent target detection model in crack identification.
[0054] The following is a detailed description of step 102, i.e., "loading the pre-trained target detection model, using the target detection model to extract features for each frame image in the continuous video image, identifying the location of the cracks through a multi-layer convolutional neural network, and generating a crack segmentation mask map; wherein, when pre-training the target detection model, extracting the optical flow features between adjacent frames in the first data set through an optical flow calculation algorithm, and deeply merging the optical flow features with the image features" in conjunction with an embodiment.
[0055] In the present invention, in order to realize real-time monitoring of bridge cracks, a target detection model based on deep learning is adopted. First, a pre-trained target detection model is loaded. The model has been trained and optimized on a pre-prepared first data set so as to accurately extract and analyze the input image data. The first data set is a key data resource for supporting the training and testing of time series prediction models. Specifically, the data set contains historical monitoring data related to crack width changes. The data comes from the long-term bridge monitoring system or on-site field measurements, reflecting the changes in crack width at different time nodes.
[0056] The target detection model of the present application can be implemented by training a variety of specific models. For example, the Faster R-CNN (Region-based Convolutional Neural Networks) model can be used. It is a deep learning model based on the region proposal network (RPN) and the convolutional neural network (CNN). It is characterized by the ability to generate candidate regions, classify these regions, and regress bounding boxes. This model performs well in terms of accuracy and is particularly suitable for tasks that require high detection accuracy. It can also be implemented using the YOLO (You Only Look Once) series of models. The YOLO series of models is a type of target detection algorithm based on deep learning. Its main feature is that the target detection task is regarded as a single regression problem, thereby achieving fast and real-time target detection. The YOLO series of models are widely used in various computer vision tasks such as video surveillance, autonomous driving, industrial detection, etc. due to their excellent detection speed and good accuracy. Preferably, the present application can use the YOLOv8 model for crack detection. YOLOv8 is the latest version of the YOLO series, which further improves performance and efficiency. It introduces some new modules and optimization techniques, such as a new backbone network structure, lightweight model design, and more efficient reasoning strategies.
[0057] When the target detection model processes each frame image in the continuous video image, the target detection model extracts features from the image through a multi-layer convolutional neural network (CNN) to identify the location of the crack. The network can automatically learn and extract key features in the image, and capture local and global information of the image at different levels through multi-layer convolution operations. The output of the model is a pixel-level segmentation mask map, which accurately marks the area of the crack in the image, facilitating subsequent analysis and width measurement. The pixel-level segmentation mask map of the crack is a binary or multi-valued image used in image processing and computer vision to represent the crack area in the image. Each pixel in this mask map has a corresponding value, indicating whether the pixel belongs to the crack area or other background area.
[0058] In the pre-training stage of the target detection model, an optical flow calculation algorithm is introduced to extract optical flow features between adjacent frames in the first data set. The optical flow algorithm is a computer vision technology used to estimate the motion of objects or scenes between video sequences or continuous image frames. The optical flow algorithm can describe the movement of each pixel in time, thereby generating a motion vector field between each frame, revealing the dynamic information of the object or scene. Commonly used optical flow algorithms include the Horn-Schunck algorithm, the Lucas-Kanade algorithm, and the Farneback algorithm, etc. This application can be flexibly selected according to needs.
[0059] In the present invention, the optical flow features between adjacent frames are calculated by the optical flow algorithm, and are deeply fused with the image features of the frame image samples. First, the motion information between adjacent frames in the frame image samples is analyzed by the optical flow calculation algorithm to obtain the dynamic change characteristics of the cracks on the time axis. Specifically, by performing pixel-level analysis on the input adjacent frame images, the motion vector of each pixel is calculated to generate an optical flow feature map, which contains the directional information and speed characteristics of the crack changes and can accurately reflect the dynamic change trend of the cracks. Secondly, the frame image samples are subjected to multi-layer convolution processing to extract the spatial features of the image and form an image feature representation. The image features contain the geometric morphology, position distribution and texture characteristics of the cracks, providing static spatial information support for the target detection model. Subsequently, the optical flow features are deeply fused with the image features, and a specific feature fusion strategy is adopted to integrate the dynamic information in the optical flow features and the static information in the image features. The fused features can not only reflect the morphological characteristics of the cracks, but also contain their dynamically changing spatiotemporal information, further improving the model's detection capability for complex cracks. Finally, the fusion features are used to optimize the training of the target detection model and adjust the model parameters so that it can effectively learn the comprehensive characteristics of cracks in spatial and temporal dimensions. In this way, the pre-trained target detection model has higher detection accuracy and robustness, providing a solid foundation for the subsequent generation of crack segmentation mask images.
[0060] As an implementable method, the present application can be implemented through the following steps when deeply fusing optical flow features with image features and normalizing the fused features before inputting them into a target detection model: calculating pixel motion vectors between consecutive frames through an optical flow algorithm and extracting inter-frame motion characteristics; performing multi-layer convolution processing on the optical flow features to generate an optical flow feature map; utilizing a channel attention mechanism to fuse the optical flow feature map with the image feature map, and retaining the key features of the original image through a residual connection; and normalizing the fused feature map before inputting it into the target detection model.
[0061] Specifically, first, the pixel motion vector between consecutive frame images is calculated by the optical flow algorithm to extract the inter-frame motion characteristics. Among them, the inter-frame motion characteristics can include the following parameters: optical flow vector, which is the core representation of inter-frame motion characteristics, defines the direction and amplitude of the movement of each pixel between two frames, and can reflect the speed and direction of the object; optical flow amplitude, which indicates the speed of pixel movement; optical flow direction, which indicates the direction of movement, expressed in radians or angles; gradient change, which reflects the change of pixel intensity in the time dimension.
[0062] Then, the extracted optical flow features are processed by multi-layer convolution to generate an optical flow feature map. Specifically, the motion vector is decomposed into a horizontal component map and vertical component graph , by calculating the magnitude of the motion vector and direction , generate additional motion characteristic channels, and integrate the above multiple channels into a multi-layer optical flow feature map. Then, the channel attention mechanism is used to deeply fuse the optical flow feature map with the image feature map. The channel attention mechanism highlights the importance of key channels by assigning different weights to different feature channels, thereby effectively retaining the motion characteristics and static image features that are crucial to crack detection during the feature fusion process. In addition, residual connections are introduced in the fusion process to ensure that key information of the original image, such as crack edges and texture features, is retained in the fused feature map. This residual connection can alleviate the problem of information loss and improve the learning efficiency and detection performance of the model. Specifically, the residual connection can be achieved by directly adding the frame image feature map to the fused high-level features, which can retain key information such as the texture and shape of the cracks in the original frame image and avoid the loss of low-level details in high-level features.
[0063] Finally, the fused feature map is normalized to ensure the consistency of the eigenvalue distribution, which is convenient for the subsequent learning and training of the target detection model. Normalization can accelerate the convergence of the model and reduce the fluctuation of model performance caused by inconsistent eigenvalue ranges. The fused feature map after normalization is input into the target detection model, and the multi-layer convolutional neural network in the model is used to perform pixel-level segmentation of cracks and generate a pixel-level segmentation mask map of the cracks. Among them, normalization can be implemented in many ways, such as Z-Score normalization, which adjusts the feature map data distribution to a standard normal distribution with a mean of 0 and a standard deviation of 1 through the mean and standard deviation; BN (Batch Normalization) can also be used to calculate the mean and standard deviation of the feature map in each batch to standardize the feature map.
[0064] As an implementable method, in the real-time monitoring method of bridge cracks based on the Internet of Things and deep learning described in this application, environmental parameter data of the bridge site is further introduced to enhance the recognition ability and detection accuracy of the target detection model for crack features. Environmental parameters can reflect the potential impact of the actual bridge environment on the formation and development of cracks. Environmental parameter data include but are not limited to temperature, humidity, wind speed and light intensity. These parameters can be obtained by installing special environmental monitoring sensors at the bridge site, or by obtaining meteorological data of the bridge area from a nearby meteorological station, or by using remote sensing equipment or environmental monitoring instruments carried by drones to regularly monitor the bridge area and collect environmental data. By jointly analyzing environmental parameters with visual information in continuous video image frames, the present invention can more comprehensively evaluate the dynamic characteristics of cracks.
[0065] Specifically, the acquired environmental parameters and video image data are simultaneously input into the target detection model. The environmental parameters are feature encoded by a multi-layer perceptron to extract their internal correlation and nonlinear features. The encoded environmental features are further deeply fused with the video image features, thereby combining environmental factors with the visual characteristics of cracks, improving the model's feature expression ability and the robustness of crack detection.
[0066] In order to achieve efficient pre-training of the target detection model, the present application expands the first data set to a data set containing environmental parameter samples. Specifically, the crack data are grouped according to the environmental parameter samples, and multiple sub-data sets are constructed so that each sub-data set can reflect the crack characteristics under different environmental conditions. For example, each sub-data set can correspond to a specific environmental condition, such as a specific temperature range, humidity range, etc. On this basis, a multi-model training strategy is adopted to train multiple sub-models separately, and the weight of each sub-model is updated according to its corresponding sub-data set. This method can significantly improve the adaptability of the model to changes in environmental conditions.
[0067] In addition, to further optimize the model performance, the weights of the feature extraction layer are shared among multiple sub-models, so that the feature extraction layer can fully learn the common features of each sub-dataset. At the same time, multiple sub-models are integrated to generate the final target detection model through a fusion strategy. The integrated target detection model not only has good versatility, but also can show high crack detection accuracy for specific environmental conditions, thereby achieving accurate detection and real-time monitoring of cracks.
[0068] The feature extraction layer usually uses a convolutional neural network or a similar architecture to extract high-level features from the input frame image. In the initial stage, the parameters of the feature extraction layer are initialized by a pre-trained model or by random initialization. The parameters of the feature extraction layer are kept consistent in multiple sub-models, and a sharing mechanism is used to ensure that all sub-models use the same feature extraction capabilities during training, thereby improving the overall performance and generalization ability of the model.
[0069] After completing the training of each sub-model, the weights of its task-specific layers are jointly optimized to ensure that the model has high detection performance under all environmental conditions. The detection results of multiple sub-models are integrated through weighted averaging or ensemble learning (such as Bagging, Boosting) to generate the final target detection model. For example: weighted voting is performed on the prediction results of each sub-model, and the output feature map of each sub-model is weighted summed to generate the final feature map. In the integration stage, a multi-task loss function can also be introduced to globally optimize the model, taking into account the detection performance under different environmental conditions.
[0070] The above step 103, i.e., "extracting the contour pixel points of the crack based on the crack segmentation mask image, calculating the pixel normal width of the crack in combination with the image gradient information, and obtaining the width change information and width change rate of the crack; the width change information includes: average width change value, local width change value and / or relative width change value," is described in detail below in conjunction with an embodiment.
[0071] First, the crack contour pixels are extracted from the crack pixel-level segmentation mask output by the target detection model. The segmentation mask clearly identifies the crack area, and the edge contour of the crack can be obtained through image processing algorithms such as edge detection (such as Canny edge detection) or contour extraction algorithms (such as the findContours method in OpenCV). This contour information is used for subsequent width analysis and transformation rate calculation.
[0072] On the basis of obtaining the crack contour information, the image gradient information is used to further analyze the crack width. By calculating the image gradient at the crack contour, the normal direction of the edge can be determined. The normal direction refers to the direction perpendicular to the contour line from the crack edge, which is used to measure the width of the crack in this direction. By sampling along the normal direction at each crack edge point and calculating the pixel spacing between samples, the pixel width of the crack can be accurately obtained.
[0073] After measuring the pixel normal width of the crack, the change in the crack width in different frame images can be analyzed to obtain the width change information of the crack. By comparing the widths of consecutive frames and combining the time interval, the change rate of the crack width is calculated, which represents the expansion speed of the crack at different time points. Among them, the width change information includes: average width change value, local width change value and / or relative width change value. The average width change value of the crack reflects the overall expansion trend of the crack; the local width change value can identify the expansion of the crack at a specific location; and the relative width change value is used to measure the change ratio of the crack at different time points or under different conditions. Figure 3 is the data of the width change information measured in the actual application of this application. As shown in Figure 3, Figure 3 (a) is a crack average width change diagram in the embodiment of this application; Figure 3 (b) is a crack local width change diagram in the embodiment of this application, specifically the average width change diagram of the first 100 pixels with the largest degree of change; Figure 3 (c) is a crack relative width change diagram in the embodiment of this application, specifically the crack average width change diagram relative to the first 5 frames.
[0074] As an implementable method, the present application can also correct the corner points in the contour, use a local curvature algorithm to calculate the curvature value of each point in the contour pixel points; identify the corner points in the contour pixel points according to the curvature value of each point; and modify the pixel points on both sides of the tangent direction of the corner point into contour pixel points through an interpolation algorithm. Specifically, first, use a local curvature algorithm to calculate the curvature value of each point in the contour pixel points. The curvature value reflects the curvature characteristics of each contour point, that is, the degree of curvature of the curve at this point. The specific implementation method includes selecting a point in a small neighborhood area and determining the curvature value by calculating the degree of curvature of the point in the area. Common curvature calculation methods can use curvature estimation based on least squares fitting, or use a curvature calculation algorithm based on difference. Secondly, identify the corner points in the contour pixel points according to the curvature value of each point. By analyzing the curvature values of each pixel point, points whose curvature values change significantly in their neighborhood are identified, which are corner points. This process can include setting a preset threshold, screening out points whose curvature values exceed the threshold as corner points, or using a local maximum detection method to automatically identify corner points. Finally, the pixel points on both sides of the tangent direction of the corner point are modified into contour pixel points through an interpolation algorithm. This step is intended to adjust the corner points to optimize the continuity and width calculation of the crack contour. In a specific implementation, a linear interpolation or polynomial interpolation method can be used to interpolate the pixel points along the tangent direction before and after the corner point, and adjust their positions to ensure that the contour line can smoothly transition at the corner point. The interpolated pixel points will be included in the definition of the crack contour, providing more accurate width analysis data and reducing the impact of errors at the corner points on subsequent calculations.
[0075] In conjunction with the embodiments below, the above step 104, namely "based on the width change information and the width transformation rate, using a time series prediction model to perform trend prediction on the crack width change to obtain a prediction result of the crack width change; wherein, the time series prediction model uses a second data set for model training, the second data set includes a second frame image sample and its corresponding prediction result sample, and during the training process of the time series prediction model, a basic model and a complexity capture module are designed, the basic model generates a first prediction result based on the input frame image sample; the complexity capture module identifies a complex area based on the input frame image sample, and generates a second prediction result for the complex area; the first prediction result and the second prediction result are weightedly fused to obtain a third prediction result; the basic model and the complexity capture module are iterated with the goal of minimizing the difference between the third prediction result and the prediction result sample; wherein, in each round of parameter iteration, the parameters of the basic model are updated using a fixed learning rate, and the parameters of the complexity capture module are updated using a dynamic learning rate" is described in detail.
[0076] In the present application, the time series prediction model performs trend prediction by analyzing crack width change information and width change rate to obtain the prediction result of crack width change. The model is trained using the second data set to ensure its accuracy and stability under different conditions.
[0077] The second data set is used for training the time series prediction model, and the data set contains historical monitoring data related to the crack width change, specifically including the second frame image sample and its corresponding prediction result sample. The second data set can select the same data sample content as the first data set as the second data set, or can select data sample content more suitable for the time series prediction model training as the second data set according to the needs.
[0078] During the training process, the model design includes two main parts: the basic model and the complexity capture module. The specific implementation method is as follows: First, the second frame image sample is used as the input of the model, and the basic model and the complexity capture module are designed to process the input data respectively. The basic model mainly models the overall change trend of the crack width. Its input is the frame image sample, and its output is the first prediction result, which reflects the global change characteristics of the crack width. Next, the complexity capture module further analyzes the input frame image samples, focusing on identifying the high-complexity areas, such as areas where the crack width changes dramatically. Through a specific complex area recognition algorithm, such as analysis based on gradient changes or local fluctuation characteristics, the module can accurately locate the complex areas of the crack, and refine these areas to model and generate the second prediction result. Subsequently, the output results of the basic model and the complexity capture module are weighted fused to generate a comprehensive third prediction result. The fusion process can adopt a weight optimization strategy, such as a weighted attention mechanism to dynamically allocate the result contributions of the basic model and the complexity capture module to improve the accuracy of the prediction results.
[0079] During the training process, the loss function (such as mean square error) is used to iteratively update the model parameters with the goal of minimizing the difference between the third prediction result and the prediction result sample. In order to improve the optimization effect, in each round of parameter iteration, the parameters of the basic model are updated with a fixed learning rate to ensure that the model can stably capture the global change characteristics of the cracks; while the parameters of the complexity capture module are updated with a dynamic learning rate to adapt to the drastic changes in complex areas, thereby enhancing the model's ability to capture detailed changes. Learning rate is a key hyperparameter in machine learning and deep learning, which is used to control the magnitude of the model's weight update in each iteration. It determines the size of the model's step size during the optimization process, thereby affecting the speed and effect of model training. By combining the outputs of the basic model and the complexity capture module, the weights of the main model are optimized in combination with a weighted fusion strategy, thereby achieving a comprehensive prediction of the crack width change trend.
[0080] For highly complex areas, this application uses the statistical characteristics of crack width changes and local features in the image to identify them. For example, the input data is first analyzed for local variance or local standard deviation to identify areas with large changes in the time series. These areas usually represent sudden changes or drastic changes in crack changes. By calculating the rate of change of the crack width and combining it with the image gradient information, the complexity of the area is further determined and marked.
[0081] During the training process of the complexity capture module, an adaptive learning rate adjustment strategy is adopted to better cope with the dynamic changes in the high-complexity areas in the data. Specifically, a gradient-based learning rate adjustment method is used to calculate the gradient amplitude of each training batch and compare it with the set threshold. When the gradient amplitude exceeds the threshold, the learning rate is dynamically increased to speed up the convergence of the model in the high-complexity area; when the gradient amplitude is small or close to zero, the learning rate is reduced to avoid overfitting of the model in the low-complexity area. This strategy ensures that the details of the high-complexity area can be fully learned and modeled in the crack width changes.
[0082] In the training of the time series prediction model, the output results of the basic model and the complexity capture module are combined through a weighted fusion strategy. Specifically, this can be achieved in the following way: first, the basic model and the complexity capture module are trained separately so that they can learn independently and produce prediction results. Then, the prediction results of each model are weighted, and the weight value is adjusted according to the performance in the validation data set. Finally, the output results of the two parts are merged using a weighted fusion strategy to form an optimized main model output. Among them, the weighted fusion strategy can be implemented in many ways. For example, if the basic model performs better in overall trend prediction, it is assigned a higher weight; and if the complexity capture module has a higher prediction accuracy in high-complexity areas, it is assigned a higher weight. This strategy can dynamically adjust the weight assignment to achieve the best prediction effect; it can also fuse the output results of the basic model and the complexity capture module through weighted summation, and continuously monitor the loss value of the model during the training process. If the loss of a model at a certain stage is significantly lower than that of other models, its weight is gradually increased, and vice versa.
[0083] As an implementable approach, the specific implementation of the basic model and complexity capture module in the time series prediction model includes: the basic model adopts a single-layer long short-term memory (LSTM) network; the complexity capture module adopts a double-layer LSTM structure to locally model the drastic changes in crack width in high-complexity areas; during the training process of the time series prediction model, a complexity evaluation function is used to dynamically adjust the learning rate of the complexity capture module according to the crack width change rate; the output results of the basic model and the complexity capture module are integrated to optimize the time series prediction model through a weighted attention mechanism.
[0084] The basic model uses a single-layer LSTM network, which can effectively capture the time series characteristics of the input data, especially showing good performance in processing the overall change trend of the crack width. The single-layer LSTM network can learn the long-term dependencies of the data and model the global trend of the crack width through a loop structure, thereby providing accurate trend prediction. The complexity capture module uses a double-layer LSTM structure, which, based on the basic model, further improves the modeling ability of local areas with drastic changes in the input data. The double-layer LSTM network effectively extracts high-order features in the time series through its multi-layer stacking characteristics, and is specifically used to capture crack width changes in high-complexity areas, such as a sharp increase or decrease in crack width.
[0085] Among them, local modeling of the drastic changes in crack width in high-complexity areas can introduce local feature extraction modules and attention mechanisms to focus specifically on the parts of the input sequence that change dramatically. Specifically, before the input of the double-layer LSTM, local high-complexity areas can be identified by calculating the crack width change rate or the complexity evaluation function based on threshold detection. These areas will receive more attention during the training process, so that the double-layer LSTM network can more effectively capture the drastic change characteristics of these areas.
[0086] Preferably, during model training, the learning rate of high complexity areas is dynamically adjusted through the complexity evaluation function. For areas with drastic changes in crack width, a higher learning rate is used to speed up the convergence of the model in these areas. For areas with smaller changes, a lower learning rate is used to avoid overfitting. The mechanism of dynamically adjusting the learning rate can make the complexity capture module more agile when dealing with drastic changes, while remaining stable when dealing with stable areas.
[0087] After completing the training of the basic model and the complexity capture module, in order to optimize the overall performance of the time series prediction model, the output results of the two modules are combined through the weighted attention mechanism. This mechanism dynamically adjusts the weights of the basic model and the complexity capture module by weighting the output of each part and combining the different characteristics of the input data. In this way, the time series prediction model can intelligently select the output of different models as the final prediction result according to the complexity of the input data and the prediction requirements, thereby improving the accuracy and robustness of the prediction.
[0088] This application evaluates the health status of the bridge through data analysis and processing based on the prediction of the crack width change trend. Specifically, according to the crack width change trend prediction results output by the time series prediction model, the crack width change rate at different time points is calculated. The risk level of bridge cracks is divided according to the predicted data of this rate. The risk level is divided according to the pre-set threshold standard. For example, according to the different crack width change rates, it is divided into three levels: low risk, medium risk and high risk.
[0089] After the assessment is completed, one or more of the bridge health assessment results, the width change information, the width change rate, and the crack width change prediction results can be transmitted through the Internet of Things platform. This platform can comprehensively analyze the real-time acquired bridge data and the prediction results, and send them to the monitoring center in the form of a data report. After receiving the data report, the monitoring center updates the bridge health status in real time, generates and displays a visual chart of the real-time health status, so that relevant personnel can monitor and analyze the condition of the bridge.
[0090] In addition, based on the prediction results of the crack width change rate, bridge maintenance recommendations are generated for different risk levels. When a high risk level is detected, crack repair, local reinforcement or comprehensive inspection measures are recommended; while at a low risk level, regular monitoring and inspection are recommended. This process ensures that the bridge maintenance strategy is consistent with the current health status, and that necessary maintenance and reinforcement measures can be taken in a timely manner based on different assessment results to ensure the safety and long-term stability of the bridge.
[0091] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0092] According to an embodiment of another aspect, a bridge crack width monitoring device is provided. Figure 4 A schematic block diagram of a bridge crack width monitoring device according to an embodiment is shown, Figure 4 As shown, the device 400 includes:
[0093] The image acquisition unit 401 is configured to acquire continuous video images of the bridge surface, and the continuous video images include multiple frame images.
[0094] The target detection unit 402 is configured to load a pre-trained target detection model, and use the target detection model to extract features from each frame image in the continuous video image to obtain image features of the frame image, and use the image features to identify the location of the crack through a multi-layer convolutional neural network to generate a segmentation mask map of the crack; wherein the target detection model uses a first data set for model training, and the first data set includes a first frame image sample and its corresponding crack identification result sample. When pre-training the target detection model, the optical flow features between adjacent frame images in the first frame image sample are extracted through an optical flow calculation algorithm, and the optical flow features are fused with the image features of the frame image sample.
[0095] The width change analysis unit 403 is configured to extract the contour pixel points of the crack based on the crack segmentation mask image, calculate the pixel normal direction width of the crack in combination with the image gradient information, and obtain the width change information and width change rate of the crack; the width change information includes: average width change value, local width change value and / or relative width change value.
[0096] The width change prediction unit 404 is configured to perform trend prediction of the crack width change based on the width change information and the width transformation rate using a time series prediction model to obtain a prediction result of the crack width change; wherein the time series prediction model uses a second data set for model training, the second data set includes a second frame image sample and its corresponding prediction result sample, and during the training process of the time series prediction model, a basic model and a complexity capture module are designed, the basic model generates a first prediction result based on the input frame image sample; the complexity capture module identifies a complex area based on the input frame image sample, and generates a second prediction result for the complex area; the first prediction result and the second prediction result are weightedly fused to obtain a third prediction result; the basic model and the complexity capture module are iterated with the goal of minimizing the difference between the third prediction result and the prediction result sample; wherein in each round of parameter iteration, the parameters of the basic model are updated using a fixed learning rate, and the parameters of the complexity capture module are updated using a dynamic learning rate.
[0097] As one of the achievable methods, the device 400 also includes an environmental information acquisition unit 405, which is configured to obtain environmental parameter data of the bridge site, and the environmental parameters include temperature, humidity, wind speed and light intensity; the target detection unit 402 is configured to: input the environmental parameters and the video image into the target detection model at the same time when using the target detection model to extract features for each frame image in the continuous video image; feature encode the environmental parameters through a multi-layer perceptron, and deeply fuse the encoded environmental features with the video image features; the first data set also includes environmental parameter samples; when pre-training the target detection model, the crack data in the first data set is grouped according to the environmental parameter samples to establish multiple sub-data sets; train multiple sub-models, and update the weight of each sub-model based on the corresponding sub-data set; share the weights of the feature extraction layer in multiple sub-models, and integrate the multiple sub-models to generate the target detection model.
[0098] As one of the feasible ways, when the target detection unit 402 deeply fuses the optical flow features with the image features and inputs the fused features into the target detection model after normalization, it can be configured as follows: calculating the pixel motion vector between consecutive frames through the optical flow algorithm and extracting the inter-frame motion characteristics; performing multi-layer convolution processing on the optical flow features to generate an optical flow feature map; utilizing the channel attention mechanism to fuse the optical flow feature map with the image features and retaining the key features of the original image through the residual connection; and normalizing the fused feature map and inputting it into the target detection model.
[0099] As one of the feasible methods, after extracting the contour pixel points of the crack based on the crack segmentation mask image, the width change analysis unit 403 is configured as follows: using a local curvature algorithm to calculate the curvature value of each point in the contour pixel points; identifying the corner points in the contour pixel points according to the curvature value of each point; and modifying the pixel points on both sides of the tangent direction of the corner point into contour pixel points through an interpolation algorithm.
[0100] As one of the feasible ways, the width change prediction unit 404 can be configured as follows when designing the basic model and the complexity capture module: the basic model adopts a single-layer long short-term memory network structure to construct the model; the complexity capture module adopts a double-layer long short-term memory network structure to locally model the drastic changes in crack width in high-complexity areas; during the training of the time series prediction model, a complexity evaluation function is used to dynamically adjust the learning rate of the complexity capture module according to the crack width change rate; the output results of the basic model and the complexity capture module are combined, and the parameters of the time series prediction model are optimized through a weighted attention mechanism.
[0101] As one of the feasible methods, after acquiring continuous video images of the bridge surface, the image acquisition unit 401 preprocesses the video images, and the preprocessing includes one or more of the following: graying the frame images; removing noise in the frame images by combining Gaussian filtering and edge-preserving filtering algorithms; improving the contrast of the frame images by contrast enhancement algorithms; and normalizing the pixel values of the frame images.
[0102] As one of the feasible ways, the device 400 also includes a result reporting unit 406, which is configured to evaluate the health status of the bridge using the predicted results of the crack width change trend, and send the width change information, the width change rate, the predicted results of the crack width change and / or the bridge health status report to the monitoring center through the Internet of Things platform. The monitoring center updates the health status assessment results in real time and displays the results visually. The health status assessment includes: based on the predicted results of the crack width change, dividing the risk level of bridge cracks, and generating bridge maintenance suggestions according to different risk levels, including crack repair, local reinforcement or comprehensive inspection.
[0103] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the unit described as a separate component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place, or it may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative work.
[0104] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0105] In addition, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored, and when the program is executed by a processor, the steps of any one of the methods in the aforementioned method embodiments are implemented.
[0106] And an electronic device, comprising: one or more processors; and a memory associated with the one or more processors, the memory being used to store program instructions, and the program instructions, when read and executed by the one or more processors, execute the steps of any of the methods described in the aforementioned method embodiments.
[0107] The present application also provides a computer program product, including a computer program, which implements the steps of any one of the methods in the aforementioned method embodiments when executed by a processor.
[0108] in, Figure 5 The architecture of the electronic device is shown as an example, which may include a processor 510, a video display adapter 511, a disk drive 512, an input / output interface 513, a network interface 514, and a memory 520. The processor 510, the video display adapter 511, the disk drive 512, the input / output interface 513, the network interface 514, and the memory 520 may be communicatively connected via a communication bus 530.
[0109] The processor 510 may be implemented by a general-purpose CPU, a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in this application.
[0110] The memory 520 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 520 can store an operating system 521 for controlling the operation of the electronic device 500, and a basic input and output system (BIOS) 522 for controlling the low-level operation of the electronic device 500. In addition, a web browser 523, a data storage management system 524, and a bridge crack width monitoring device 525, etc. can also be stored. The above-mentioned bridge crack width monitoring device 525 can be an application program that specifically implements the operations of the aforementioned steps in the embodiment of the present application. In short, when the technical solution provided by the present application is implemented by software or firmware, the relevant program code is stored in the memory 520 and is called and executed by the processor 510.
[0111] The input / output interface 513 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0112] The network interface 514 is used to connect to a communication module (not shown) to achieve communication interaction between the device and other devices. The communication module can achieve communication through a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WIFI, Bluetooth, etc.).
[0113] The bus 530 comprises a pathway for transmitting information between the various components of the device (eg, the processor 510 , the video display adapter 511 , the disk drive 512 , the input / output interface 513 , the network interface 514 , and the memory 520 ).
[0114] It should be noted that, although the above device only shows a processor 510, a video display adapter 511, a disk drive 512, an input / output interface 513, a network interface 514, a memory 520, a bus 530, etc., in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include components necessary for implementing the solution of the present application, and does not necessarily include all the components shown in the figure.
[0115] It can be known from the description of the above implementation methods that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application can be essentially or partly contributed to the prior art in the form of a computer program product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application or certain parts of the embodiments.
[0116] The technical solution provided by the present application is described in detail above. The principle and implementation method of the present application are described in detail using specific examples. The description of the above embodiments is only used to help understand the method and core idea of the present application. At the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as limiting the present application.
Claims
1. A method for monitoring bridge crack width, characterized in that: The method comprises: Acquire continuous video images of the bridge surface, wherein the continuous video images include a plurality of frame images; The target detection model obtained by pre-training is loaded, and the target detection model is used to extract features of each frame image in the continuous video image to obtain image features of the frame image, and the image features are used to identify the position of the crack through a multi-layer convolutional neural network to generate a crack segmentation mask map; wherein the target detection model is trained using a first data set, and the first data set includes a first frame image sample and its corresponding crack identification result sample. When pre-training the target detection model, an optical flow feature between adjacent frame images in the first frame image sample is extracted through an optical flow calculation algorithm, and the optical flow feature is fused with the image feature of the frame image sample; Extracting the outline pixel points of the crack based on the crack segmentation mask image, calculating the pixel normal direction width of the crack in combination with the image gradient information, and obtaining the width change information and width change rate of the crack; the width change information includes: average width change value, local width change value and / or relative width change value; According to the width change information and the width transformation rate, a time series prediction model is used to predict the trend of the crack width change to obtain a prediction result of the crack width change; wherein the time series prediction model takes the frame image sample as the main input, and performs joint prediction in combination with the width change information and the width transformation rate as auxiliary parameters; the time series prediction model uses a second data set for model training, and the second data set includes a second frame image sample and its corresponding prediction result sample; during the training process of the time series prediction model, a basic model and a complexity capture module are designed; the basic model generates a first prediction result reflecting the overall crack width change trend based on the input frame image sample; the complexity capture module identifies complex areas based on the input frame image sample, and generates a locally refined second prediction result for the complex area; the first prediction result and the second prediction result are weightedly fused to obtain a comprehensive third prediction result; the basic model and the complexity capture module are iterated with the goal of minimizing the difference between the third prediction result and the prediction result sample; wherein in each round of parameter iteration, the parameters of the basic model are updated using a fixed learning rate, and the parameters of the complexity capture module are updated using a dynamic learning rate.
2. The method according to claim 1, characterized in that The method further includes: acquiring environmental parameter data at the bridge site, the environmental parameters including temperature, humidity, wind speed and light intensity; The method of extracting features from each frame image in the continuous video image by using the target detection model comprises: inputting the environmental parameters and the video image into the target detection model at the same time; encoding the environmental parameters by a multi-layer perceptron, and deeply fusing the encoded environmental features with the video image features; The first data set also includes environmental parameter samples; during pre-training, the target detection model groups the crack data in the first data set according to the environmental parameter samples to establish multiple sub-data sets; Train multiple sub-models, and update the weight of each sub-model based on the corresponding sub-dataset; The weights of the feature extraction layer are shared among the multiple sub-models, and the multiple sub-models are integrated to generate the target detection model.
3. According to the method of claim 1, the fusing of the optical flow features with the image features of the frame image samples comprises: The pixel motion vector between consecutive frames is calculated by the optical flow algorithm to extract the inter-frame motion characteristics; Performing multi-layer convolution processing on the optical flow features to generate an optical flow feature map; The optical flow feature map is fused with the image feature by using a channel attention mechanism, and the key features of the original image are retained by a residual connection; The fused feature map is normalized and then input into the target detection model.
4. The method according to claim 1, after extracting the contour pixel points of the crack based on the crack segmentation mask image, further comprises: Calculate the curvature value of each point in the contour pixel by using a local curvature algorithm; According to the curvature value of each point, identifying the corner points in the contour pixel points; By using an interpolation algorithm, the pixel points on both sides of the tangent direction of the corner point are modified into contour pixel points.
5. The method according to claim 1, characterized in that The design base model and complexity capture module include: The basic model uses a single-layer long short-term memory network structure to build the model; The complexity capture module adopts a two-layer long short-term memory network structure to locally model the dramatic changes in crack width in high-complexity areas; During the training of the time series prediction model, a complexity evaluation function is used to dynamically adjust the learning rate of the complexity capture module according to the crack width change rate; The output results of the basic model and the complexity capture module are integrated, and the parameters of the time series prediction model are optimized through a weighted attention mechanism.
6. The method according to claim 1, further comprising: After acquiring continuous video images of the bridge surface, the video images are preprocessed, and the preprocessing includes one or more of the following: Performing grayscale processing on the frame image; A Gaussian filter and an edge-preserving filter algorithm are used to remove noise in the frame image; Using a contrast enhancement algorithm to improve the contrast of the frame image; The pixel values of the frame image are normalized.
7. The method according to claim 1, further comprising: The health status of the bridge is evaluated by using the prediction result of the crack width change trend, and the width change information, the width change rate, the prediction result of the crack width change and / or the bridge health status report are sent to the monitoring center through the Internet of Things platform. The monitoring center updates the health status evaluation result in real time and displays the result visually; The health status assessment includes: dividing the risk level of bridge cracks based on the predicted results of the crack width change, and generating bridge maintenance suggestions according to different risk levels, including crack repair, local reinforcement or comprehensive inspection.
8. A bridge crack width monitoring device, characterized in that: The device comprises: An image acquisition unit is configured to acquire continuous video images of a bridge surface, wherein the continuous video images include a plurality of frame images; The target detection unit is configured to load a pre-trained target detection model, use the target detection model to perform feature extraction on each frame image in the continuous video image, obtain image features of the frame image, use the image features to identify the location of the crack through a multi-layer convolutional neural network, and generate a crack segmentation mask map; wherein the target detection model uses a first data set for model training, the first data set includes a first frame image sample and its corresponding crack identification result sample, when pre-training the target detection model, extracts optical flow features between adjacent frame images in the first frame image sample through an optical flow calculation algorithm, and fuses the optical flow features with the image features of the frame image sample; A width change analysis unit is configured to extract the contour pixel points of the crack based on the crack segmentation mask image, calculate the pixel normal direction width of the crack in combination with the image gradient information, and obtain the width change information and width change rate of the crack; the width change information includes: an average width change value, a local width change value and / or a relative width change value; The width change prediction unit is configured to use a time series prediction model to perform trend prediction on the crack width change according to the width change information and the width change rate, and obtain a prediction result of the crack width change; wherein the time series prediction model uses the frame image sample as the main input, and performs joint prediction in combination with the width change information and the width change rate as auxiliary parameters; the time series prediction model uses a second data set for model training, and the second data set includes a second frame image sample and its corresponding prediction result sample; during the training process of the time series prediction model, a basic model and a complexity capture module are designed, and the basic model generates a first prediction result reflecting the overall crack width change trend based on the input frame image sample; the complexity capture module performs complex area recognition based on the input frame image sample, and generates a locally refined second prediction result for the complex area; the first prediction result and the second prediction result are weightedly fused to obtain a comprehensive third prediction result; the basic model and the complexity capture module are iterated with the goal of minimizing the difference between the third prediction result and the prediction result sample; wherein in each round of parameter iteration, the parameters of the basic model are updated using a fixed learning rate, and the parameters of the complexity capture module are updated using a dynamic learning rate.
9. An electronic device, characterized in that: include: one or more processors; And a memory associated with the one or more processors, the memory is used to store program instructions, and when the program instructions are read and executed by the one or more processors, the steps of the method described in any one of claims 1 to 7 are executed.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Crack monitoring device for sea-crossing bridge and monitoring method thereof
CN118730914A
Municipal pipeline management method and system based on positioning sensing function
CN119067466A