Underwater bridge pier crack detection method, system, equipment, medium and product
The improved CycleGAN and Deeplabv3+ networks underwater pier crack detection solve the problems of low detection efficiency and insufficient accuracy, and achieve efficient and accurate identification of underwater pier cracks.
Patent Information
- Application Number
- CN202510585402.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The prior art has problems such as low detection efficiency, insufficient accuracy, high cost, and limited deep learning model training and generalization capabilities in underwater bridge pier crack detection, especially in complex environments, it is difficult to effectively identify diversified cracks.
Data enhancement is used to enhance data, combined with the improved Deeplabv3+ network, and feature enhancement is used to enhance feature enhancement through the encoder, CBAM attention mechanism module and the decoder's dimensionality-update module, and the surface cracks of the underwater piers are identified, and the dual attention mechanism module is used to enhance feature enhancement of multi-scale feature maps.
The accuracy and efficiency of underwater piers crack detection are improved, the interference of the problem of few samples is solved, and the adaptability and detection performance of the model are enhanced.
Smart Images

Figure CN120495235A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and in particular to a method, system, equipment, medium and product for detecting cracks in underwater bridge piers. Background Art
[0002] As a crucial component of cross-river and cross-sea bridges, bridge piers are prone to cracking under extreme loads and the combined effects of multiple factors. Underwater crack detection faces numerous challenges, including complex environments, diverse crack forms, and difficulty identifying them. Traditional manual inspections are time-consuming, inefficient, and have a high rate of missed detections. Sensor monitoring is costly, and sonar and radar have limited detection depths, preventing comprehensive inspections of deepwater areas. They also suffer from large positioning errors and poor adaptability. Using underwater robots to detect cracks on bridge pier surfaces is a future trend in intelligent emergency rescue equipment. Underwater robots play a vital role in underwater engineering construction and maintenance, including dam and tunnel foundation inspections, submarine cable and pipeline maintenance, and offshore platform structure monitoring and repair.
[0003] Underwater crack detection technology is of great significance in promoting the intelligent application of robots. However, the scarcity of real underwater crack datasets limits the training and generalization capabilities of advanced technology models such as deep learning, thus hindering the further application and development of the technology. Summary of the Invention
[0004] The purpose of this application is to provide a method, system, equipment, medium and product for underwater pier crack detection, which can improve the accuracy of underwater pier crack detection based on an improved CycleGAN network and an improved Deeplabv3+ network.
[0005] To achieve the above objectives, this application provides the following solutions:
[0006] In a first aspect, the present application provides a method for detecting cracks in underwater bridge piers, comprising:
[0007] Real-time acquisition of underwater bridge pier images taken by underwater robots;
[0008] The trained improved CycleGAN network is used to perform data enhancement on the underwater bridge pier image to obtain an enhanced image; the improved CycleGAN network includes an encoder, a CBAM attention mechanism module and a decoder; a dimension increase-reduction module is provided between the encoder and the decoder;
[0009] The enhanced image is segmented using a trained improved Deeplabv3+ network to identify surface cracks on underwater bridge piers; the improved Deeplabv3+ network includes a dual attention mechanism module, which is used to perform feature enhancement on multi-scale feature maps.
[0010] In a second aspect, the present application provides an underwater bridge pier crack detection system, comprising:
[0011] An image acquisition module is used to acquire underwater bridge pier images taken by an underwater robot in real time;
[0012] A data enhancement module is used to perform data enhancement on the underwater bridge pier image using the trained improved CycleGAN network to obtain an enhanced image; the improved CycleGAN network includes an encoder, a CBAM attention mechanism module and a decoder; a dimensionality increase-reduction module is provided between the encoder and the decoder;
[0013] A crack recognition module is used to segment the enhanced image using a trained improved Deeplabv3+ network to identify cracks on the surface of underwater bridge piers; the improved Deeplabv3+ network includes a dual attention mechanism module, which is used to perform feature enhancement on multi-scale feature maps.
[0014] In a third aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned underwater pier crack detection method.
[0015] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the above-mentioned underwater pier crack detection method when executed by a processor.
[0016] In a fifth aspect, the present application provides a computer program product, including a computer program, which implements the above-mentioned underwater pier crack detection method when executed by a processor.
[0017] According to the specific embodiments provided in this application, this application has the following technical effects:
[0018] This application combines an improved CycleGAN network with an improved DeepLabV3+ network, leveraging the domain adaptability of the CycleGAN network in few-shot learning and the lightweight and high-precision advantages of the DeepLabV3+ network. This effectively addresses the interference caused by the few-shot problem in underwater bridge pier crack detection and improves the accuracy of underwater bridge pier crack detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0020] Figure 1 This is a diagram of an application environment of a method for detecting cracks in underwater piers in one embodiment of the present application;
[0021] Figure 2 A schematic flow chart of a method for detecting cracks in underwater bridge piers provided in one embodiment of the present application;
[0022] Figure 3 Schematic diagram of the structure of the improved CycleGAN network;
[0023] Figure 4 This is a schematic diagram of the structure of the CBAM attention mechanism module;
[0024] Figure 5 This is a schematic diagram of the structure of the improved Deeplabv3+ network;
[0025] Figure 6 Schematic diagram of the structure of the efficient channel attention module;
[0026] Figure 7 Schematic diagram of the loss function curve during improved Deeplabv3+ network training;
[0027] Figure 8 Schematic diagram of the results of image segmentation using different algorithms;
[0028] Figure 9 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0029] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0030] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0031] The underwater pier crack detection method provided in the embodiment of the present application can be applied to Figure 1In the application environment shown, the terminal 102 communicates with the server 104 via a network. The data storage system can store data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send the underwater pier image taken by the underwater robot to the server 104. After the server 104 receives the underwater pier image taken by the underwater robot, the server 104 uses the trained improved CycleGAN network to perform data enhancement on the underwater pier image to obtain an enhanced image, and uses the trained improved Deeplabv3+ network to segment the enhanced image and identify cracks on the surface of the underwater pier. The server 104 can feedback the obtained underwater pier surface cracks to the terminal 102. In addition, in some embodiments, the underwater pier crack detection method can also be implemented independently by the server 104 or the terminal 102. For example, the terminal 102 can directly perform crack detection on the underwater pier images taken by the underwater robot, or the server 104 can obtain the underwater pier images taken by the underwater robot from the data storage system and perform crack detection on the underwater pier images taken by the underwater robot.
[0032] Terminal 102 may include, but is not limited to, various desktop computers, laptops, smartphones, tablet computers, IoT devices, and portable wearable devices. IoT devices may include smart speakers, smart TVs, smart air conditioners, and smart car devices. Portable wearable devices may include smart watches, smart bracelets, and head-mounted devices. Server 104 may be implemented as a standalone server or a server cluster consisting of multiple servers, or may be a cloud server.
[0033] In an exemplary embodiment, Figure 2 As shown, a method for detecting cracks in underwater piers is provided. The method is executed by a computer device, specifically a computer device such as a terminal or a server, or a terminal and a server. In the embodiment of the present application, the method is applied to Figure 1 The server 104 in the example is used to illustrate the process, which includes the following steps S1 to S3.
[0034] S1: Real-time acquisition of underwater bridge pier images taken by underwater robots.
[0035] In this embodiment, an underwater robot is used to capture images of underwater bridge piers. The underwater robot mainly consists of a main frame, an electronic sealed cabin, a propulsion device, a deep learning image processing module, and a lighting system.
[0036] The electronic sealed cabin, propulsion system, lighting system, deep learning image processing module, etc. are all arranged on the main frame, and the corresponding structures are fixed to the main frame with bolts. Considering that the underwater robot needs to meet the requirements of rigidity and strength as well as possess certain corrosion resistance and ductility, the main frame is made of high-molecular nylon material. At the same time, four buoyancy materials are installed above the main frame for balancing. The buoyancy material also has certain corrosion resistance and ductility.
[0037] The electronic sealed cabin is mainly used to place various drive control boards and detection equipment to ensure that these electronic components will not be damaged by seawater pressure and seawater corrosion. At the same time, the electronic sealed cabin occupies a part of the mass of the underwater robot. In this implementation, the underwater robot will adopt a cylindrical design with a hemispherical head. The hemispherical head is connected to the cabin body with an O-ring seal. The other end of the cabin body is provided with waterproof docking sockets for various components and waterproof sockets for communication cables. The cylindrical cabin body of the electronic sealed cabin is made of aluminum alloy, and the hemispherical head adopts an acrylic hemispherical cover. A large number of key electronic components are arranged inside the electronic sealed cabin, such as the underlying hardware driver board, power management system board, high-performance lithium battery and other electronic components;
[0038] In addition, in order to have a wider observation range, a 1080P pitch underwater camera is placed in the hemispherical observation cabin of the sealed cabin, and there is enough space to set the camera's pitch and other degrees of freedom; this camera is used to more conveniently observe the environment inside the aquarium, and at the same time, in order to expand the observation range, an external 1080P underwater camera is placed to observe the surface conditions of underwater bridge piers.
[0039] The propulsion device mainly consists of 6 propeller thrusters and an underlying hardware driver board. The underwater robot realizes multi-posture control of the underwater robot by cooperatively controlling the operation of the 6 propeller thrusters to meet the requirements of stability and flexibility of the underwater robot's image acquisition. Among them, 2 propeller thrusters are used to realize the suspension control of the underwater robot in the vertical direction, and the other 4 propeller thrusters are arranged in the horizontal direction to realize propulsion in all directions. In this embodiment, the propeller thrusters use an airfoil-shaped planing slurry with better fluid effect. The airfoil-shaped planing slurry will help to obtain better propulsion force. The propeller thrusters will be driven by the underlying hardware driver board, and each propeller drive signal line will be connected to the hardware driver board. Each propeller drive power line will be connected to the power management system board. The power management system board can not only realize the effective identification of the battery power, but also has the functions of balanced charging and discharging, temperature protection, MOS switch, hardware voltage stabilization power supply, etc.
[0040] The deep learning image processing module primarily consists of a camera acquisition device and a deep learning image processing control board. To mitigate the impact of weak underwater light on image recognition, this embodiment selects a starlight-level low-light camera as the main acquisition device and is equipped with a lighting system. The camera data cable will be connected to the signal input terminal of the deep learning image processing control board. To expand the robot's detection field of view, the underwater robot will use a two-degree-of-freedom camera gimbal for large-scale searches. The camera and the two-degree-of-freedom gimbal are connected with bolts to enable pitch and yaw movement of the camera module. The camera will capture the visual image in front of the underwater robot and transmit this visual image information to the deep learning image processing control board for image processing. The deep learning image processing control board integrates and deploys an improved CycleGAN network and an improved Deeplabv3+ network to process the captured image information in real time and transmit the processed image information back for display through the output terminal of the deep learning image processing control board.
[0041] The lighting system primarily consists of two high-power underwater lights. Good lighting conditions are crucial for underwater observation. Water has poor light transmittance, and light attenuates rapidly in water. To meet the functional requirements of observing objects in low-light environments, the underwater robot is equipped with an auxiliary lighting system. The lights utilize P70HI-LED bulbs, housed in a 30° focused, pressure-resistant glass lampshade. The entire housing is CNC-machined from 6061 aluminum alloy, offering corrosion and pressure resistance. The lights are evenly spaced on either side of the main frame.
[0042] S2: Using the trained improved CycleGAN network to perform data enhancement on the underwater pier image to obtain an enhanced image. Figure 3 As shown, the improved CycleGAN network includes an encoder, a CBAM (Convolutional Block Attention Module) attention mechanism module, and a decoder. A dimensionality increase / reduction module is provided between the encoder and the decoder. The dimensionality increase / reduction module includes a residual connection module and a dimensionality reduction convolution kernel. The CBAM attention mechanism module includes a channel attention module and a spatial attention module.
[0043] This embodiment leverages the feature alignment capabilities of the improved CycleGAN network's encoder to effectively incorporate the style of the real underwater environment used for training into the crack image data to be detected. This process aligns the ground concrete crack images used for training with the underwater crack images to be detected in feature space, minimizing potential interference from background differences on crack feature extraction.
[0044] The specific working process of the improved CycleGAN network is as follows:
[0045] like Figure 3 As shown in the figure, after the underwater bridge pier image passes through the encoder composed of convolutions to extract features, a residual connection mechanism (implemented by the residual connection module) is used to transmit some feature information directly to the decoder output. This mechanism effectively reduces the information loss caused by increasing network depth. After the convolution operation, batch normalization is uniformly applied to improve training stability, followed by a ReLU activation function to introduce nonlinear characteristics.
[0046] Before the input features enter the decoding stage, the features of the encoder output and the residual connection are fused in the channel dimension. Part of the feature fusion uses 1×1 convolution for channel compression and information mapping. Figure 3 The arrows in the middle indicate the direction of feature flow and inter-layer connections. This enhances information exchange between different features and improves feature expression capabilities. Through the above optimization, this embodiment effectively reduces feature loss during the transmission of input feature information, improving the generation accuracy of the improved CycleGAN network, making it suitable for image style transfer, medical image enhancement, and other image generation tasks that require preserving detailed features.
[0047] In order to make the encoder of the CycleGAN network pay more attention to key feature information, this embodiment introduces the CBAM attention mechanism module in the feature extraction residual block part. The structure of the enhanced residual block is as follows Figure 4 shown.
[0048] In order to improve the feature extraction capability, this application designs an improved residual block. The input features are processed through three different paths and then spliced in the channel dimension to form rich output features. Among them, the first path passes through the channel attention module and the spatial attention module in sequence to highlight the key channels and spatial position information; the second path is used to extract local fine-grained features and enhance feature expression through padding, convolution and batch normalization operations; the third path directly passes the input features to the output, retaining the original information to ensure gradient stability. Finally, the features of the three paths are spliced together to integrate information at different levels, thereby improving the model's perception and expression capabilities of crack detail features.
[0049] The CBAM attention mechanism module, based on a feature optimization method based on channel and spatial attention mechanisms, aims to improve the CycleGAN network's focus on important features, thereby enhancing the network's expressiveness and recognition accuracy. The CBAM attention mechanism module achieves dynamic weighted feature processing through the combined action of the channel and spatial attention modules.
[0050] After input features enter the channel attention module, global average pooling and global max pooling are first performed on the feature map of each channel to extract the overall channel information. The pooling results are then processed by a shared multi-layer perceptron (MLP) to learn the weight relationship between channels. The channel attention weight is generated using the sigmoid activation function. This weight is used to reweight the input channel features to highlight key feature information.
[0051] After channel weighting, the input features enter the spatial attention module. During this process, global max pooling and global average pooling are first performed on the input features channel by channel, and the resulting results are concatenated along the channel dimension. Next, convolution is used to reduce the dimensionality to a single-channel feature, and a sigmoid function is applied to generate spatial attention weights. Finally, these weights are element-wise multiplied with the input features to enhance the representation of important regions in the spatial dimension.
[0052] In summary, to further optimize feature transfer, this embodiment introduces the CBAM attention mechanism module, enabling effective superposition of input information as it passes through the convolutional layers, residual connections, and attention modules. This design not only maintains the characteristics of global connectivity but also reduces information loss through residual connections, preventing network performance degradation and thus improving the model's stability and generalization capabilities.
[0053] S3: The enhanced image is segmented using the trained improved Deeplabv3+ network to identify cracks on the underwater bridge pier surface. The improved Deeplabv3+ network includes a dual attention mechanism module, which is used to enhance the multi-scale feature map. The dual attention mechanism module includes a CBAM attention mechanism module and an Efficient Channel Attention (ECA) module. The CBAM attention mechanism module in the improved Deeplabv3+ network has the same structure as the CBAM attention mechanism module in the improved CycleGAN network.
[0054] This embodiment uses the trained improved Deeplabv3+ network to accurately segment the underwater bridge pier image enhanced by the improved CycleGAN network. In traditional neural networks, continuous convolution processing of images often leads to the loss of image information. In particular, crack information usually only occupies a relatively small part in the image, which makes it particularly difficult to retain and extract crack features. To solve this problem, this embodiment designs a feature enhancement structure that integrates multiple attention mechanisms based on the original Deeplabv3+ network. The structure of the improved Deeplabv3+ network is as follows: Figure 5 shown
[0055] The specific working process of the improved Deeplabv3+ network is as follows:
[0056] In this embodiment, the enhanced image is a color crack image of size 512. After being processed by the MobilenetV2 backbone network, feature maps of different scales are obtained. Among them, the feature map of size 64 is enhanced by the CBAM attention mechanism module to increase the focus on key information. At the same time, the feature map of size 32 is introduced into the ECA module to enhance the feature representation capability between channels. In the feature fusion stage, first, the feature map of size 32 is expanded by an upsampling convolution operation, converted into a feature map of size 64, and spliced with the 64-size feature map processed by the CBAM attention mechanism module in the channel dimension to enhance the information expression capability. Subsequently, the above-mentioned fused feature map is further upsampled to expand it to size 128 so that it can be integrated with the feature map output by another branch. Finally, the feature map expanded to size 128 is spliced with the 128-size feature map after processing and upsampling by the void space pyramid pooling in the channel dimension to achieve efficient fusion of multi-scale features.
[0057] like Figure 6 As shown in the figure, the ECA module balances the capture of local and global information by adaptively selecting the size of the convolution kernel, and the size of the convolution kernel is dynamically adjusted according to the number of channels in the input feature map. First, the input feature map is input into the global average pooling layer, and the features in each channel are average pooled in the spatial dimension to obtain a channel description vector with the same dimension as the number of channels in the input feature map. Subsequently, the channel description vector is input into the one-dimensional convolution layer for processing, and the local interaction relationship between channels is modeled through the convolution operation of the local receptive field, thereby generating the weight coefficient corresponding to each channel. The one-dimensional convolution operation uses a fixed-size convolution kernel to reduce computational complexity and avoids the large amount of parameter overhead brought by the traditional fully connected network. Furthermore, the output vector after the one-dimensional convolution processing is input into the Sigmoid activation function to normalize the weight coefficient of each channel so that the weight value of each channel falls between 0 and 1. Finally, the normalized channel weight coefficient is multiplied channel by channel with the original input feature map according to the channel dimension to achieve channel recalibration of the original feature map, so as to enhance the feature channels with higher weights and suppress the feature channels with lower weights, and finally output the feature map after attention weighting.
[0058] In this embodiment, the training samples of the improved CycleGAN network and the improved Deeplabv3+ network include ground crack images, algorithm-generated underwater crack images, and real underwater crack images.
[0059] The training sample contains a total of 9,954 images and 9,954 labels. These images are divided into training, validation, and test sets. The 4,932 images of ground cracks, each sized 448×448, are divided into an 8:2 ratio. The training set contains 3,945 images, and the validation set contains 987 images. Furthermore, the improved CycleGAN network is used to perform style transfer on the 4,932 ground crack images, resulting in 4,932 underwater crack images, which are also divided into the same ratio. The remaining 90 images, each sized 512×512 (real underwater crack images), serve as the test set.
[0060] The training set guides the optimization of model parameters, and the model uses this data to adjust its internal weights. The validation set is used for evaluation during or after training. The test set is used to evaluate the model's final performance, focusing on verifying its generalization capabilities and simulating unknown data that the model may encounter in real-world scenarios.
[0061] The loss functions of the improved Deeplabv3+ network during training include the cross entropy loss function and the Dice loss function.
[0062] Cross Entropy Loss Function CE loss is a standard way to evaluate the difference between the predicted probability distribution and the true label distribution in classification tasks:
[0063]
[0064] y i is the true label of the current category, p(y i ) is the probability of the current category, and N is the number of categories.
[0065] In image segmentation tasks, the cross-entropy loss function directly measures the difference between the probability distribution of the model output and the true label, providing a clear goal for model optimization. Since this function directly incorporates the classification result of each pixel into the loss calculation, it effectively promotes model optimization in terms of improving the accuracy of pixel-level predictions. However, cross-entropy loss calculates the average loss across all categories. When the number of samples of certain categories in a dataset is significantly higher than that of other categories, the model tends to optimize those categories that appear more frequently. This phenomenon occurs because, from a probabilistic perspective, improving the prediction accuracy of common categories leads to a more significant reduction in the overall loss, making the model less sensitive to categories with fewer samples. In our dataset, since the crack region accounts for a small proportion of the entire image, relying solely on cross-entropy loss may lead to severe class imbalance.
[0066] The Dice loss function is an indicator that measures the similarity between two sets and is often used to compare the predicted results with the actual results in image segmentation tasks:
[0067]
[0068] X represents the pixel labels of the true segmented image, while Y represents the pixel categories predicted by the model for the segmented image. The numerator is approximately the dot product between the predicted image pixels and the true label image pixels, that is, the sum of the corresponding pixel products. The denominator represents the sum of the pixels in their respective corresponding images. Therefore, the Dice loss focuses on the overlapping area between the prediction and the true label, rather than pure pixel accuracy. This allows the model to naturally focus on fewer positive samples (cracks) during training, thereby reducing the excessive influence of the majority class (background). As a result, the Dice loss is less affected by class imbalance. However, the Dice loss may not be sensitive enough to small object regions or discontinuous object boundaries. When the object region is very small or highly dispersed, even a small prediction deviation can lead to a significant change in the Dice coefficient. This can cause the model to under-learn these small regions, affecting the overall segmentation quality.
[0069] To sum up, the loss function Loss of the Deeplabv3+ network during training is:
[0070] Loss=CE loss +Dice loss
[0071] The combination of Dice loss and cross entropy loss in the loss function enables the simultaneous optimization of global accuracy and local overlap in image segmentation tasks, improves training stability and overcomes the challenge of class imbalance. This design not only stabilizes the training process, but also improves the overall performance of the model in terms of precision, recall, and specificity through their respective advantages. It is particularly suitable for crack image segmentation tasks that require precise contours. Figure 7 As can be seen from the figure, the loss function of the model eventually converges to a value between 0.1 and 0.15. This result verifies the effectiveness of the generated dataset and the effectiveness of the training of the improved Deeplabv3+ network.
[0072] To validate the detection effectiveness of the improved Deeplabv3+ network, comparative experiments were conducted between the network before and after the improvements. This work consists of three parts. First, training loss curves were compared to verify the improved Deeplabv3+ network's improved feature extraction capabilities during training. Second, experiments were designed and validated using evaluation metrics for the proposed improvements. Finally, images of underwater bridge pier surface defect detection were presented, visualizing the effectiveness of the improved Deeplabv3+ network.
[0073] First, to verify the effectiveness of the improved loss function, experimental results of the improved Deeplabv3+ network's loss function were compared with those of the original Deeplabv3+ network trained using only the cross-entropy loss function. This confirmed that combining cross-entropy and Dice loss effectively mitigates the impact of class imbalance. Next, a dual-attention mechanism was introduced into the structure of the original Deeplabv3+ network and compared with the original model. Finally, further comparative experiments were conducted using the improved CycleGAN and the Deeplabv3+ network.
[0074] The ablation test results of the model on the validation set are shown in Table 1. Comparing the results of Experiments 1 and 2, we show that by combining the cross-entropy loss with the Dice loss function (i.e., the CE-dice loss), the recall rate and intersection-over-union ratio increased by 2.35% and 1.48%, respectively. This is because the traditional cross-entropy loss function provides stability during training but has limitations when dealing with class imbalance. The Dice loss function, on the other hand, can alleviate the impact of class imbalance to a certain extent by assigning higher weights to minority class objects and lowering the weights of majority class objects. Therefore, the design of the CE-Dice loss function helps improve the detection performance of crack objects.
[0075] In comparative experiments 1 and 3, the original Deeplabv3+ network structure was optimized and an attention mechanism was introduced to further enhance the feature fusion structure. This optimization improved the model's recall and intersection-over-union (IoU) by 0.82% and 1.48%, respectively. The optimized network structure not only preserves the original information of small objects in the image but also strengthens the model's ability to extract this information.
[0076] Comparative Experiment 1 and Experiment 4 show that after combining loss function optimization and model structure improvement, the recall rate and intersection-over-union ratio increased by 2.83% and 1.82% respectively. Finally, in Comparative Experiment 1 and Experiment 5, by combining loss function optimization, network structure improvement and the introduction of Cycle-Deeplabv3+ architecture, the recall rate and intersection-over-union ratio increased by 3.93% and 2.22% respectively, achieving the best results. This is because the Cycle-Deeplabv3+ architecture can unify the background style of the data set while retaining the characteristics of the cracks, thereby reducing the interference of the background on the crack characteristics. The final results prove that this application significantly improves the detection performance by unifying the style of the data set while strengthening small target detection, effectively verifying the effectiveness of the model improvement.
[0077] Table 1
[0078]
[0079] The results of the model on the test set are shown in Table 2. Comparing Experiments 1 and 2, optimizing the loss function improved recall and IoU by 0.32% and 0.31%, respectively. Comparing Experiments 1 and 3, improving the model structure improved recall and IoU by 5.73% and 6.21%. Comparing Experiments 1 and 4, the combined optimization of the loss function and the improvements to the model structure resulted in improvements of 11.18% and 11.89%, respectively.
[0080] Table 2
[0081]
[0082] However, the performance of the first four experiments on the test set was significantly lower than that on the validation set. This was primarily due to the significant differences in style between the real fracture images and the training images. Although fracture features shared some similarity, the model was unable to effectively identify fracture features in the real underwater fracture images due to limitations in the dataset and the model's inherent performance, as well as the interference of complex backgrounds.
[0083] To further verify the generalization performance of the model, we compared it with a traditional segmentation algorithm. The experimental results are shown in Table 3. From the results in Table 3, we can see that this application achieves a good balance between detection performance and lightweightness.
[0084] Table 3
[0085]
[0086] After the model training is completed, a set of images will be randomly collected to show the visualization effect of the detection results of this application and the above-mentioned traditional segmentation algorithm, such as Figure 8 shown. Figure 8 The results show that Unet and Segformer, as algorithms with the most complex network structures, significantly outperform Pspnet and Hrnet in detection performance. However, for more complex underwater crack image recognition, these algorithms perform poorly in crack feature extraction. In contrast, the visualization results of the GAN-Deeplab algorithm proposed in this paper can more accurately extract crack features from underwater crack images in complex backgrounds, demonstrating its advantages in processing such images.
[0087] Based on the same inventive concept, embodiments of the present application also provide an underwater bridge pier crack detection system. The implementation solution provided by this system is similar to the implementation solution described in the aforementioned method. Therefore, the specific limitations in one or more embodiments of the underwater bridge pier crack detection system provided below can be found in the aforementioned limitations of the underwater bridge pier crack detection method and will not be further elaborated here.
[0088] In an exemplary embodiment, a system for detecting cracks in an underwater bridge pier is provided, comprising:
[0089] The image acquisition module is used to acquire the underwater bridge pier images taken by the underwater robot in real time.
[0090] A data enhancement module is used to perform data enhancement on the underwater bridge pier image using a trained improved CycleGAN network to obtain an enhanced image; the improved CycleGAN network includes an encoder, a CBAM attention mechanism module and a decoder; a dimensionality increase-reduction module is provided between the encoder and the decoder.
[0091] A crack recognition module is used to segment the enhanced image using a trained improved Deeplabv3+ network to identify cracks on the surface of underwater bridge piers; the improved Deeplabv3+ network includes a dual attention mechanism module, which is used to perform feature enhancement on multi-scale feature maps.
[0092] In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above-mentioned method embodiments. The computer device can be a server or a terminal, and its internal structure can be as shown in FIG. Figure 9 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store data to be processed. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for detecting cracks in underwater bridge piers is implemented.
[0093] Those skilled in the art will understand that Figure 9The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application and does not constitute a limitation on the computer device to which the solution of the present application is applied. A specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps of the above-mentioned method embodiments when executing the computer program.
[0094] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0095] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0096] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0097] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0098] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.
[0099] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0100] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A method for detecting cracks in underwater bridge piers, characterized in that: include: Real-time acquisition of underwater bridge pier images taken by underwater robots; The trained improved CycleGAN network is used to perform data enhancement on the underwater bridge pier image to obtain an enhanced image; the improved CycleGAN network includes an encoder, a CBAM attention mechanism module and a decoder; a dimension increase-reduction module is provided between the encoder and the decoder; The enhanced image is segmented using a trained improved Deeplabv3+ network to identify surface cracks on underwater bridge piers; the improved Deeplabv3+ network includes a dual attention mechanism module, which is used to perform feature enhancement on multi-scale feature maps.
2. The underwater pier crack detection method according to claim 1, characterized in that: The dimension increase-dimension reduction module includes a residual connection module and a dimension reduction convolution kernel.
3. The underwater pier crack detection method according to claim 1, characterized in that: The dual attention mechanism module includes a CBAM attention mechanism module and an efficient channel attention module.
4. The underwater pier crack detection method according to claim 1 or 3, characterized in that: The CBAM attention mechanism module includes a channel attention module and a spatial attention module.
5. The underwater pier crack detection method according to claim 3, characterized in that: The convolution kernel size of the efficient channel attention module is dynamically adjusted according to the number of channels in the input feature map.
6. The underwater pier crack detection method according to claim 1, characterized in that: The loss function Loss of the improved Deeplabv3+ network during training is: Loss=CE loss +Dice loss Among them, CE loss is the cross entropy loss function, Dice loss is the Dice loss function, y i is the true label of the current category, p(y i ) is the probability of the current category, X is the pixel label of the real segmented image, N is the number of categories, and Y is the pixel category predicted by the improved Deeplabv3+ network for the segmented image.
7. An underwater pier crack detection system, characterized in that: include: An image acquisition module is used to acquire underwater bridge pier images taken by an underwater robot in real time; A data enhancement module is used to perform data enhancement on the underwater bridge pier image using the trained improved CycleGAN network to obtain an enhanced image; the improved CycleGAN network includes an encoder, a CBAM attention mechanism module and a decoder; a dimensionality increase-reduction module is provided between the encoder and the decoder; A crack recognition module is used to segment the enhanced image using a trained improved Deeplabv3+ network to identify cracks on the surface of underwater bridge piers; the improved Deeplabv3+ network includes a dual attention mechanism module, which is used to perform feature enhancement on multi-scale feature maps.
8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the underwater pier crack detection method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the underwater pier crack detection method according to any one of claims 1 to 6 is implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the underwater pier crack detection method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Cross-sea bridge pier surface crack detection method based on linear residual attention
CN116012310A
Mask-based water conveyance tunnel underwater crack image enhancement method
CN117575967A
Water tunnel crack detection method based on UNet network
CN117952898A
Water tunnel crack detection method based on YOLOv4 network
CN118864347A
Imaging logging image crack segmentation method based on domain adaptation
CN118898711A