Hydraulic structure crack detection method, device, electronic equipment and storage medium

Through the layered multi-resolution feature aggregation module and depth separation convolution technology, combined with the composite attention mechanism and activation function, the problem of insufficient accuracy in traditional image crack detection technology in complex environments is solved, and high-precision, real-time and anti-interference crack detection effect is achieved.

CN119722677BActive Publication Date: 2025-05-06SICHUAN ENERGY INTERNET RES INST TSINGHUA UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510228476.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-06
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

Traditional image crack detection technology has the problem of insufficient crack inspection accuracy in the context of complex background environment, uneven light and subtle and diverse cracks.

Method used

The layered multi-resolution feature aggregation module is used for feature fusion, combined with the depth separation convolution and composite attention mechanism, the image is extracted and weighted, and finally feature enhancement is performed through the activation function to obtain crack detection results.

Benefits of technology

In a complex background, the accuracy and robustness of crack detection are significantly improved, the significance of crack characteristics is enhanced, real-time detection on edge devices is supported, and anti-interference ability is strong.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119722677B_ABST
    Figure CN119722677B_ABST
Patent Text Reader

Abstract

The present invention provides a method, device, electronic device and storage medium for detecting cracks in hydraulic structures, and relates to the field of image detection technology. The method comprises: obtaining an image to be detected of a hydraulic structure, performing rough segmentation on the image to be detected, and obtaining a rough segmentation image; performing feature aggregation on the image to be detected and the rough segmentation image, and obtaining a feature fusion image; performing feature extraction on the feature fusion image by using a deep separable convolution, and obtaining a feature extraction image; performing weighting on the feature extraction image by using a composite attention mechanism, and obtaining a feature weighted image; performing feature enhancement on the feature weighted image by using an activation function, and obtaining a crack detection result. The present invention performs well in crack detection tasks, not only significantly improving detection accuracy, but also having strong real-time and anti-interference capabilities, and providing strong technical support for health monitoring of hydraulic structures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image detection, and in particular to a method, device, electronic equipment and storage medium for detecting cracks in a hydraulic structure. Background Art

[0002] As key infrastructure in water conservancy projects, hydraulic structures such as dams and cushion ponds are of vital importance for their safe and stable operation. However, various factors such as concrete material degradation, load action, uneven foundation settlement, and long-term fatigue service can easily lead to cracks in hydraulic structures. These cracks may not only cause structural problems, but also accelerate steel corrosion, damage the concrete protective layer, further expand cracks, and seriously affect the durability of hydraulic structures. Therefore, it is necessary to detect cracks in hydraulic structures.

[0003] Traditional manual crack detection methods have many shortcomings, such as large workload, low efficiency, and are easily affected by the subjective judgment of the inspectors. Especially for complex crack types such as shear cracks, they are prone to missed detection and misjudgment, and can no longer meet the needs of fast and accurate detection of large-scale structures. In recent years, automatic crack detection technology based on computer vision has received extensive attention and research due to its advantages such as high precision, fast analysis and automated processing. However, in the case of complex background environment, uneven lighting, and subtle and diverse cracks, traditional image crack detection technology still has the problem of insufficient crack inspection accuracy. Summary of the invention

[0004] In view of the above-mentioned deficiencies in the prior art, the present invention provides a hydraulic structure crack detection method, device, electronic equipment and storage medium, which effectively solve the problem of insufficient crack inspection accuracy.

[0005] In a first aspect, the present invention provides a method for detecting cracks in a hydraulic structure, the method comprising:

[0006] Acquire an image of a hydraulic structure to be detected, and roughly segment the image to be detected to obtain a roughly segmented image;

[0007] Performing feature aggregation on the image to be detected and the roughly segmented image to obtain a feature fused image;

[0008] Using depthwise separable convolution to perform feature extraction on the feature fusion image to obtain a feature extraction image;

[0009] Using a composite attention mechanism to weight the feature extraction image to obtain a feature weighted image;

[0010] An activation function is used to enhance the features of the feature weighted image to obtain a crack detection result.

[0011] Furthermore, the expression of the feature aggregation is as follows:

[0012]

[0013] In the above formula, F agg represents the feature fusion result, represents the number of layers of the input feature map, L represents the total number of layers of the input feature map, F l Indicates The input feature map of the layer, Conv 1×1 represents a 1×1 convolution operation, u l Represents the operation of upsampling input feature maps of different resolutions to the target size.

[0014] Furthermore, the extracting features from the feature fused image using depthwise separable convolution includes:

[0015] A deep convolution is performed on each channel of the feature fusion image to obtain a deep convolution result. The formula of the deep convolution is as follows:

[0016]

[0017] In the above formula, Represents the output result of deep convolution, i and j Represent the row index and column index on the feature map respectively, c represents the depthwise convolution channel, m and n Respectively represent the width and height of the depth convolution kernel, K Indicates the total amount of offset, Indicates the depth of the convolution kernel at the offset ( m , n ) and in the channel c The weight on

[0018] The depth convolution result is convolved point by point, and the formula of the point by point convolution is as follows:

[0019]

[0020] In the above formula, Represents the point-by-point convolution output result, k represents the point-wise convolution channel, Represents the depth convolution channel c With point-wise convolution channels k The weight between C Indicates the total number of channels.

[0021] Furthermore, the composite attention mechanism is used to weight the feature extraction image, including:

[0022] Using channel attention and spatial attention to weight the feature extraction image in the spatial dimension to obtain an initial weighted feature map;

[0023] Performing global average pooling on the initial weighted feature map and generating an adaptive weight vector;

[0024] The channel importance of the initial weighted feature map is weighted according to the adaptive weight vector.

[0025] Furthermore, the adopting of channel attention and spatial attention to weight the feature extraction image in the spatial dimension includes:

[0026] The feature extraction image is weighted using channel attention, and the expression for weighting the channel attention is as follows:

[0027]

[0028] In the above formula, M c represents the channel attention weight, Represents the Sigmoid function, Re LU represents a nonlinear activation function, AVGPool ( F ) represents the global maximum pooling of the input feature map. MaxPool ( F ) represents global average pooling of the input feature map. W c1 and W c2 Respectively represent the weight matrix;

[0029] The feature extraction image is weighted using spatial attention, and the expression for weighting the spatial attention is as follows:

[0030]

[0031] In the above formula, M s represents the spatial attention weight, Conv 7×7 Represents a 7×7 convolution operation.

[0032] Further, weighting the channel importance of the initial weighted feature map according to the adaptive weight vector includes:

[0033] Channel attention is used to weight the importance of the channel. The expression for weighting the channel attention is as follows:

[0034]

[0035] In the above formula, M se represents the channel importance weight, Represents the Sigmoid function, Re LU represents a nonlinear activation function, represents the adaptive weight vector, W se1 and W se2 They represent weight matrices respectively.

[0036] Furthermore, the step of using an activation function to enhance the features of the feature weighted image includes:

[0037] The Swish activation function is used to enhance the crack edge details in the feature weighted image to obtain an initial detection image;

[0038] The Mish activation function is used to converge the crack edge features in the initial detection image.

[0039] In a second aspect, the present invention provides a hydraulic structure crack detection device, the device comprising:

[0040] An image processing module is used to obtain an image of a hydraulic structure to be detected, and to roughly segment the image to be detected to obtain a roughly segmented image;

[0041] A feature aggregation module, used for performing feature aggregation on the image to be detected and the coarse segmentation image to obtain a feature fusion image;

[0042] A feature extraction module, used to extract features from the feature fusion image using depthwise separable convolution to obtain a feature extraction image;

[0043] A feature weighting module, used for weighting the feature extraction image by using a composite attention mechanism to obtain a feature weighted image;

[0044] The crack detection module is used to enhance the features of the feature weighted image by using an activation function to obtain a crack detection result.

[0045] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the hydraulic structure crack detection method as described in the first aspect of the present invention.

[0046] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the hydraulic structure crack detection method as described in the first aspect of the present invention.

[0047] The hydraulic structure crack detection method, device, electronic device and storage medium provided by the present invention introduce a hierarchical multi-resolution feature aggregation module to achieve effective fusion of features of different scales to meet the detection needs of various crack sizes. By calculating an efficient deep separable convolution structure, the computational complexity is reduced, thereby supporting real-time detection on edge devices. The composite attention mechanism can focus on the crack area more accurately under complex backgrounds, effectively suppress irrelevant features in the background, and enhance the significance of crack features, while maintaining high detection accuracy and robustness under complex backgrounds. It performs well in crack detection tasks, not only significantly improving detection accuracy, but also having strong real-time and anti-interference capabilities, providing strong technical support for health monitoring of hydraulic structures. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0049] Figure 1 It is a schematic diagram of the flow of a hydraulic structure crack detection method provided by an embodiment of the present invention;

[0050] Figure 2 is a schematic diagram of the architecture of a depthwise separable convolution in an embodiment of the present invention;

[0051] Figure 3 is a schematic diagram of the structure of the convolutional block attention module in an embodiment of the present invention;

[0052] Figure 4 is a schematic diagram of the structure of the SE attention module in an embodiment of the present invention;

[0053] Figure 5 is a schematic diagram of a network structure of a crack detection model in an embodiment of the present invention;

[0054] Figure 6 is a schematic diagram of a training loss value in an embodiment of the present invention;

[0055] Figure 7 is a schematic diagram of crack detection results of a crack detection model in an embodiment of the present invention;

[0056] Figure 8is a schematic diagram comparing crack detection results of various models in an embodiment of the present invention;

[0057] Fig. 9 It is a structural schematic diagram of a hydraulic structure crack detection device provided by an embodiment of the present invention;

[0058] Fig.10 It is a structural schematic diagram of an electronic device provided by an embodiment of the present invention.

[0059] Description of main component symbols:

[0060] 200, hydraulic structure crack detection device; 210, image processing module; 220, feature aggregation module; 230, feature extraction module; 240, feature weighting module; 250, crack detection module; 300, electronic device; 310, processor; 320, communication interface; 330, memory; 340, communication bus. DETAILED DESCRIPTION

[0061] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be further clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. It should be noted that the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0062] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0064] Crack detection in hydraulic structures is extremely important. Traditional artificial crack detection methods have many shortcomings, such as large workload, low efficiency, and susceptibility to subjective judgment by inspectors. Especially for complex crack types such as shear cracks, missed detection and misjudgment are prone to occur, which makes it difficult to meet the needs of fast and accurate detection of large-scale structures. In recent years, automatic crack detection technology based on computer vision has received extensive attention and research due to its advantages such as high precision, fast analysis and automated processing. However, in the case of complex background environment, uneven lighting, and subtle and diverse cracks, traditional image crack detection technology still has the problem of insufficient crack inspection accuracy.

[0065] Example 1

[0066] The embodiment of the present invention provides a hydraulic structure crack detection method, which effectively solves the problem of insufficient crack detection accuracy in traditional image crack detection technology. Figure 1 FIG. 1 is a flow chart of a method for detecting cracks in hydraulic structures provided by an embodiment of the present invention. Figure 1 As shown, the method includes:

[0067] S100, obtaining an image of a hydraulic structure to be detected, and roughly segmenting the image to be detected to obtain a roughly segmented image.

[0068] In the embodiment of the present invention, the hydraulic structure includes but is not limited to structural facilities such as dams and water cushion ponds, and the image to be detected is an image of the part of the hydraulic structure that needs to be crack detected. After image preprocessing, edge detection, image segmentation, morphological operation, connected region analysis and contour extraction are performed on the image to be detected, a coarse segmentation image is obtained. The coarse segmentation of the image to be detected can quickly identify the crack area from the image, reduce the amount of data for subsequent processing, and improve the processing speed and accuracy.

[0069] S200 , performing feature aggregation on the image to be detected and the coarsely segmented image to obtain a feature fused image.

[0070] In the embodiment of the present invention, since the sizes and shapes of cracks in hydraulic structures such as dams and water cushion ponds vary, it is necessary to effectively capture crack features at multiple scales. The embodiment of the present invention uses a hierarchical multi-resolution feature aggregation module to extract and fuse feature maps at different resolutions to obtain a feature fusion image, thereby balancing global context information and local details, and has a strong adaptability in crack detection at different scales. Among them, the low-resolution feature map provides rich context information, while the high-resolution feature map retains important detail information.

[0071] The expression for feature aggregation in the hierarchical multi-resolution feature aggregation module is as follows:

[0072]

[0073] In the above formula, F agg represents the feature fusion result, l represents the number of layers of the input feature map, L represents the total number of layers of the input feature map, F l Indicates The input feature map of the layer, Conv 1×1 represents a 1×1 convolution operation, u l Represents the operation of upsampling input feature maps of different resolutions to the target size.

[0074] Through multi-scale feature aggregation, cracks can be effectively perceived at multiple scales, and feature fusion images can be obtained, making feature expression richer and more comprehensive. In addition, when crack morphology and size vary greatly, not only the overall shape of large-scale structures can be perceived, but also tiny crack details can be captured. This multi-scale perception capability improves the performance of the detection method in different crack morphologies, enabling it to adapt to application requirements in various complex environments in dams and water cushion ponds, and maintain high detection accuracy.

[0075] S300, using depthwise separable convolution to perform feature extraction on the feature fusion image to obtain a feature extraction image.

[0076] In actual crack detection, real-time performance is an important consideration, especially in the scenarios of dams and water cushion ponds, which usually need to be deployed on embedded devices or edge devices. In this embodiment of the present invention, a deep separable convolution is used to extract features from the feature fusion image. Figure 2 is a schematic diagram of the architecture of the depthwise separable convolution in an embodiment of the present invention, such as Figure 2 As shown, the depthwise separable convolution decomposes the standard convolution into depthwise convolution and pointwise convolution.

[0077] In the embodiment of the present invention, firstly, a deep convolution is performed on each channel of the feature fusion image to obtain a deep convolution result. The formula of the deep convolution is as follows:

[0078]

[0079] In the above formula, Represents the output result of deep convolution, i and j Represent the row index and column index on the feature map respectively, c represents the depthwise convolution channel, m and n Respectively represent the width and height of the depth convolution kernel, K Indicates the total amount of offset, Indicates the depth of the convolution kernel at the offset ( m , n ) and in the depth convolution channel c The weight on .

[0080] Then perform point-by-point convolution on the depth convolution result to achieve channel fusion. The formula for point-by-point convolution is as follows:

[0081]

[0082] In the above formula, Represents the point-by-point convolution output result, k represents the point-wise convolution channel, Represents the depth convolution channel c With point-wise convolution channels k The weight between C Indicates the total number of channels.

[0083] In this embodiment of the present invention, if there is an input feature map , H represents the height of the input feature map, W represents the width of the input feature map, C Represents the number of channels of the input feature map. Through the decomposition of the depthwise separable convolution, the computational complexity of the convolution is reduced from Reduce to ,in K 2 represents the size of the convolution kernel, Represents the number of output channels of point-by-point convolution, which significantly reduces the computational complexity and memory requirements of the detection method, allowing the detection method to run more efficiently while maintaining high detection accuracy.

[0084] S400, using a composite attention mechanism to weight the feature extraction image to obtain a feature weighted image.

[0085] The crack detection task in dams and water cushion ponds is usually affected by complex backgrounds, such as concrete textures, rust stains, and wet stains. These complex background interferences may cause the model to misdetect or miss detection. In an embodiment of the present invention, a composite attention mechanism is used to weight the feature extraction image. The composite attention mechanism combines the convolutional block attention module and the SE (Squeeze-and-Excitation) attention module to focus on the crack area from the spatial and channel dimensions.

[0086] Figure 3 is a schematic diagram of the structure of the convolutional block attention module in an embodiment of the present invention. Figure 3 As shown, the convolutional block attention module weights the feature extraction image through channel attention and spatial attention, so that it can enhance the attention to cracks.

[0087] Channel attention is used to weight the feature extraction image. The expression of channel attention weighting is as follows:

[0088]

[0089] In the above formula, M c represents the channel attention weight, Represents the Sigmoid function. In order to ensure that the channel attention weight is between 0 and 1, the Sigmoid function is used to calculate the attention weight on the channel. LU represents a nonlinear activation function, AVGPool ( F ) represents the global maximum pooling of the input feature map. MaxPool ( F ) represents global average pooling of the input feature map. W c1 and W c2 They respectively represent weight matrices, which can be set according to actual conditions.

[0090] Spatial attention is used to weight the feature extraction image. Spatial attention is weighted in the spatial dimension so that the model can locate the crack area more accurately. The expression for weighting spatial attention is as follows:

[0091]

[0092] In the above formula, M s represents the spatial attention weight, Conv 7×7 Represents a 7×7 convolution operation.

[0093] The convolutional block attention module multiplies the output features of the channel attention unit and the spatial attention unit element-wise to obtain the final initial weighted feature map.

[0094] Figure 4 is a schematic diagram of the structure of the SE attention module in an embodiment of the present invention. Figure 4 As shown, x represents the feature extraction image, , and Represent feature extraction images xThe number of channels, height and width of U represent the initial weighted feature map, C, H and W represent the number of channels, height and width of the initial weighted feature map U respectively, and X represents the feature weighted image. The SE attention module further enhances the saliency of cracks through channel attention. First, the initial weighted feature map is globally averaged and pooled, and an adaptive weight vector is generated to weight the channel importance. Channel attention is used to weight the channel importance. The expression for weighting channel attention is as follows:

[0095]

[0096] In the above formula, M se represents the channel importance weight, Represents the Sigmoid function, Re LU represents a nonlinear activation function, represents the adaptive weight vector, W se1 and W se2 They respectively represent weight matrices, which can also be set according to actual conditions.

[0097] Through this channel-level weighted operation, the channel features of the crack area can be refined, so that the crack features are enhanced in the channel dimension and a feature-weighted image is obtained.

[0098] By integrating the spatial and channel attention mechanisms, the composite attention mechanism can effectively focus on the crack area in a complex background. It has better anti-interference ability and can effectively suppress irrelevant features in the background while improving the significance of crack features, thus maintaining high detection accuracy and robustness in complex backgrounds.

[0099] S500, using an activation function to enhance the features of the feature weighted image to obtain crack detection results.

[0100] In crack detection, cracks usually have complex edges and irregular shapes, and the traditional ReLU activation function has limitations in nonlinear expression capabilities. In the embodiment of the present invention, the nonlinear enhancement mechanism of the Swish activation function and the Mish activation function is introduced to achieve higher feature expression capabilities, thereby performing better in capturing crack edges and detail features. This smooth nonlinear activation function improves the overall flexibility and edge feature expression capabilities of the method, enabling it to better handle complex crack structures in dams and water cushion ponds.

[0101] The Swish activation function is used to enhance the crack edge details in the feature-weighted image. The Swish activation function enhances the expression ability of the activation function by introducing an adaptive scaling factor. The formula of the Swish activation function is as follows:

[0102]

[0103] In the above formula, represents the Sigmoid function, Indicates an adjustable parameter.

[0104] The smoothness and nonlinearity of the Swish activation function make the gradient flow smoother and help better preserve edge details to obtain the initial detection image.

[0105] The Mish activation function is used to converge the crack edge features in the initial detection image. The Mish activation function has better smoothness than the Swish activation function. The expression of the Mish activation function is as follows:

[0106]

[0107] In the above formula, e Represents the base of natural logarithms.

[0108] The Mish activation function provides a smoother gradient in the negative region, which can achieve more stable convergence on complex features such as crack edges to obtain crack detection images.

[0109] The Swish activation function and the Mish activation function enhance the nonlinear feature expression ability of the detection method, and are more precise in detecting the edges and irregular areas of cracks. Through these smooth activation functions, the subtle features of the cracks can be better captured, the detection accuracy of small cracks and crack edges can be improved, and ultimately the overall detection effect can be improved.

[0110] As a preferred implementation of an embodiment of the present invention, a crack detection model is constructed based on a hierarchical multi-resolution feature aggregation module, a depth-separable convolution, a composite attention mechanism and an activation function. Figure 5 is a schematic diagram of the network structure of the crack detection model in an embodiment of the present invention. Figure 5 As shown in the figure, the backbone network of the crack detection model adopts a combination of MobileNetV2 and VGG16 to make full use of its lightweight design and multi-scale feature expression capabilities, providing a solid basic feature for the crack detection task.

[0111] MobileNetV2, with its lightweight architecture and deep separable convolution design, can effectively reduce the number of model parameters and computational complexity, making the model run efficiently on edge devices. In addition, MobileNetV2's reverse residual structure and linear bottleneck design further enhance the ability of feature extraction, ensuring that the model has powerful representation capabilities at low computational overhead.

[0112] The shallow convolutional features of VGG16 are used to supplement multi-scale information. Its wider convolutional channels and stacked convolutional layers can extract rich low-level and mid-level features, enabling the model to capture detailed information. This feature extraction method is particularly suitable for crack detection tasks, because crack morphology is often irregular and requires features at different levels to accurately locate the crack position.

[0113] By integrating the high efficiency of MobileNetV2 and the multi-scale feature expression of VGG16, the backbone network of the crack detection model in the embodiment of the present invention can provide sufficient and effective feature representation for subsequent feature aggregation and attention modules, thereby improving the performance of the crack detection model in complex crack detection scenarios.

[0114] In order to verify the effectiveness of the hydraulic structure crack detection method provided by the embodiment of the present invention, a test experiment was conducted on the crack detection model in the embodiment of the present invention, wherein the data processing involved in the embodiment of the present invention can be developed on the RTX4090 graphics workstation, and the experimental data used is obtained by an image acquisition system designed with a robot arm, a high-definition industrial camera, a structured light camera, and a fill light source. The robot arm and the end detection module are the core equipment, which mainly complete the internal crack image acquisition of the hydraulic structure. The fill light source takes into account the weak light inside the bridge tower to ensure that the crack image with high imaging quality can be acquired. The binocular structured light camera mainly obtains its defect depth image and ranging capability.

[0115] Due to the high resolution of the original images collected by high-definition industrial cameras and the minimum depth distance imaging limitation of binocular structured light, the embodiment of the present invention crops the large-size original images into image blocks of 1024×1024 pixels through a cropping operation, sets the original data in a ratio of 9:1 between the training set and the validation set to generate 828 concrete crack photos for the training set and 92 concrete crack photos for the validation set, and uses the open source software imgAnnotation for manual annotation.

[0116] To ensure the training effect, the total number of training rounds is 100 rounds, and the learning rate is 0.0001. To ensure the uniformity of the training process, the number of batches in the training is 4. The loss function used in the process adopts a combined loss function. By combining the cross entropy loss function and the Dice loss function, the problem of some pixel prediction errors on unbalanced data in the model is alleviated, which leads to training difficulties. The cross entropy loss function is often used for classification problems. It appears to compensate for the saturation of the derivative form of the sigmoid function, thereby avoiding gradient diffusion. Figure 6 FIG. 1 is a schematic diagram of the training loss value of the crack detection model in an embodiment of the present invention. The training set loss value, the validation set loss value and the convergence of various indicators during the training process of the crack detection model are shown in FIG. Figure 6 shown.

[0117] Since the proportion of crack pixels in the concrete defect image in the embodiment of the present invention is too small and there is a serious sample imbalance problem, in order to ensure the rationality of the indicators, the model evaluation indicators in the embodiment of the present invention are mainly for defective pixels, and the model evaluation indicators include but are not limited to total pixel accuracy, recall, defect pixel precision, intersection over union (IoU), and harmonic mean F1-score, etc. In the embodiment of the present invention, the total pixel accuracy, recall, intersection over union, and harmonic mean reached 70.68%, 83.99%, 61.52%, and 76.76%, respectively.

[0118] Figure 7 FIG. 1 is a schematic diagram of crack detection results of a crack detection model according to an embodiment of the present invention. Figure 7 As shown in the figure, the crack detection model can capture the detailed features of the cracks very well. Among them, the crack detection model performs particularly well in the crack edges and connection areas. The segmentation results have strong continuity and clear boundaries, and can accurately segment irregular and smaller crack features. In addition, the crack detection model has excellent robustness to cracks in complex backgrounds, and can effectively eliminate the interference of background textures and noise, reducing false detections. Compared with the discontinuities, fuzzy edges or false detection areas that are common in the segmentation results of other models, the output results of the crack detection model of the embodiment of the present invention show curves and structures that are closer to the actual crack morphology, ensuring the integrity and accuracy of the crack area.

[0119] In order to further verify the effectiveness of each module of the crack detection model in the embodiment of the present invention, four groups of ablation experiments are designed to compare the effects of different module combinations on the performance of the crack detection model. The specific results are shown in Table 1.

[0120] Table 1. Comparison results of indicators of different module combinations

[0121]

[0122] Combination 1 includes a composite attention mechanism and a hierarchical multi-resolution feature aggregation module. It enhances the saliency of the crack area through the attention mechanism and captures cracks of different sizes using multi-scale feature fusion. Experimental results show that this combination has good anti-interference ability and multi-scale crack detection effect under complex backgrounds.

[0123] Combination 2 includes a hierarchical multi-resolution feature aggregation module and Swish and Mish activation functions. Multi-scale feature aggregation ensures the adaptability of the model to wide and narrow cracks, while Swish and Mish activation functions enhance the segmentation of fine cracks and edge details. Experimental results show that this combination performs well in detail segmentation, but its anti-interference ability is slightly insufficient.

[0124] Combination 3 includes a composite attention mechanism and Swish and Mish activation functions. The composite attention mechanism module improves the attention of the crack area, while the Swish and Mish activation functions perform well in the fine segmentation of the crack edges. This combination performs well in anti-interference and edge details, but its multi-scale crack recognition capability is slightly lacking.

[0125] Combination 4 includes all modules, namely the crack detection model in the embodiment of the present invention. Experimental results show that the crack detection model is optimal in terms of robustness under complex backgrounds, recognition of multi-scale cracks and segmentation of edge details, verifying the effectiveness of the synergy of each module and the superiority of the crack detection model.

[0126] In order to further verify the superiority of the hydraulic structure crack detection method provided by the embodiment of the present invention, under the same configuration and experimental parameters, the crack detection model in the crack detection method in the embodiment of the present invention is compared with several existing classic network models DeepCrack model, DenseNet model, Enet model, ResNet model and UNet model. The experiment is based on a self-made concrete cushion pond and dam crack data set, which covers a variety of crack characteristics, including complex background, multi-scale cracks and irregular crack shapes. Through a unified training and testing process, the comparability of the results of each model is guaranteed. The experimental results are shown in Table 2.

[0127] Table 2. Schematic diagram of the comparison results of indicators of each model

[0128]

[0129] The crack detection model of the embodiment of the present invention shows better comprehensive performance in detection accuracy, robustness and efficiency, and is more suitable for the actual application requirements of crack detection.

[0130] In order to more intuitively demonstrate the segmentation effect of the model, the embodiment of the present invention performs segmentation reasoning on the test set of the dam crack data set, and uses the total pixel precision, recall rate, intersection-over-union ratio, harmonic mean and overall accuracy as evaluation indicators to visualize the detection results of the crack detection model of the embodiment of the present invention and the comparison network models DeepCrack model, DenseNet model, Enet model, ResNet model and UNet model.

[0131] Figure 8 is a schematic diagram comparing the crack detection results of each model in the embodiment of the present invention, such as Figure 8 As shown, when comparing the detection results of each model with the real label, the performance differences of the crack detection model in terms of missed detection and false detection can be more clearly observed. The DeepCrack model has some performance in the continuity of slender cracks, but there are still a small number of false detections under noise interference. The DenseNet model uses its dense connection structure to achieve good results in the extraction of small cracks, but it is easy to lose details at the edges of the cracks, resulting in partial discontinuity in the segmentation results. As a lightweight model, the Enet model can achieve fast reasoning, but due to the limitations of the network structure, it has obvious deficiencies in segmentation accuracy and detail restoration, and the false detection phenomenon is more significant. The ResNet model has a strong feature expression ability when detecting cracks of complex shapes, but its deeper network structure leads to missed detection on small-scale cracks. The UNet model performs well in the restoration of crack edge details through a symmetrical encoding-decoding structure, but there are still false detections and noise residues in some areas. The crack detection model of the embodiment of the present invention shows obvious continuity in the output segmentation results, with clear crack edges and complete structures. Compared with other network models, there is no obvious missed detection and false detection, and it is closer to the real crack morphology. Therefore, the crack detection model of the embodiment of the present invention performs well in comprehensive segmentation effect and can extract dam crack features more accurately and robustly.

[0132] The hydraulic structure crack detection method provided by the embodiment of the present invention introduces a hierarchical multi-resolution feature aggregation module to achieve effective fusion of features of different scales to meet the detection requirements of various crack sizes. By using a computationally efficient deep separable convolution structure, the computational complexity is reduced to support real-time detection on edge devices. The composite attention mechanism can focus on the crack area more accurately under complex backgrounds, effectively suppress irrelevant features in the background, and enhance the significance of crack features, thereby maintaining high detection accuracy and robustness under complex backgrounds.

[0133] Example 2

[0134] Based on the same technical concept as the above method embodiment 1, the embodiment of the present invention provides a hydraulic structure crack detection device, Fig. 9is a schematic diagram of the structure of a hydraulic structure crack detection device provided by an embodiment of the present invention, such as Fig. 9 As shown, the hydraulic structure crack detection device 200 includes:

[0135] The image processing module 210 is used to obtain the image to be detected of the hydraulic structure, and to roughly segment the image to be detected to obtain a roughly segmented image.

[0136] The feature aggregation module 220 is used to perform feature aggregation on the image to be detected and the coarse segmentation image to obtain a feature fusion image.

[0137] The feature extraction module 230 is used to extract features from the feature fusion image using depthwise separable convolution to obtain a feature extracted image.

[0138] The feature weighting module 240 is used to weight the feature extraction image using a composite attention mechanism to obtain a feature weighted image.

[0139] The crack detection module 250 is used to enhance the features of the feature weighted image using an activation function to obtain a crack detection result.

[0140] The hydraulic structure crack detection device provided in the embodiment of the present invention performs well in crack detection tasks. It not only has a significant improvement in detection accuracy, but also has strong real-time and anti-interference capabilities, providing strong technical support for the health monitoring of hydraulic structures.

[0141] It can be understood that the implementation method of the hydraulic structure crack detection method described in the above embodiment 1 is also applicable to this embodiment and can achieve the same technical effect, so it will not be repeated here.

[0142] Example 3

[0143] Based on the same concept, an embodiment of the present invention further provides an electronic device, Fig.10 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention, such as Fig.10 As shown, the electronic device 300 may include: a processor 310, a communication interface 320, a memory 330 and a communication bus 340, wherein the processor 310, the communication interface 320, and the memory 330 communicate with each other through the communication bus 340. The processor 310 may call the logic instructions in the memory 330 to execute the steps of the hydraulic structure crack detection method described in the above embodiments. For example, it includes:

[0144] S100, obtaining an image of a hydraulic structure to be inspected, and roughly segmenting the image to be inspected to obtain a roughly segmented image;

[0145] S200, performing feature aggregation on the image to be detected and the roughly segmented image to obtain a feature fused image;

[0146] S300, extracting features from the feature fusion image using depthwise separable convolution to obtain a feature extraction image;

[0147] S400, using a composite attention mechanism to weight the feature extraction image to obtain a feature weighted image;

[0148] S500, using an activation function to enhance the features of the feature weighted image to obtain crack detection results.

[0149] The processor 310 may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or a combination of the above chips.

[0150] In addition, the logic instructions in the above-mentioned memory 330 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0151] The memory 330 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the processor, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0152] Example 4

[0153] Based on the same concept, an embodiment of the present invention further provides a computer-readable storage medium, which stores a computer program, and the computer program includes at least one code segment, which can be executed by a main control device to control the main control device to implement the steps of the hydraulic structure crack detection method described in the above embodiments. For example, it includes:

[0154] S100, obtaining an image of a hydraulic structure to be inspected, and roughly segmenting the image to be inspected to obtain a roughly segmented image;

[0155] S200, performing feature aggregation on the image to be detected and the roughly segmented image to obtain a feature fused image;

[0156] S300, extracting features from the feature fusion image using depthwise separable convolution to obtain a feature extraction image;

[0157] S400, using a composite attention mechanism to weight the feature extraction image to obtain a feature weighted image;

[0158] S500, using an activation function to enhance the features of the feature weighted image to obtain crack detection results.

[0159] Based on the same technical concept, an embodiment of the present invention further provides a computer program, which is used to implement the above method embodiment when executed by a main control device.

[0160] The computer program may be stored in whole or in part on a computer-readable storage medium packaged with the processor, or may be stored in whole or in part on a memory not packaged with the processor.

[0161] Based on the same technical concept, an embodiment of the present invention further provides a processor, which is used to implement the above method embodiment. The above processor can be a chip.

[0162] In summary, the hydraulic structure crack detection method, device, electronic device and storage medium provided by the present invention introduce a hierarchical multi-resolution feature aggregation module, which realizes the effective fusion of features of different scales to meet the detection needs of various crack sizes. By calculating the efficient depth-separable convolution structure, the computational complexity is reduced, thereby supporting real-time detection on edge devices. The composite attention mechanism can focus on the crack area more accurately under complex backgrounds, effectively suppress irrelevant features in the background, and enhance the significance of crack features, while maintaining high detection accuracy and robustness under complex backgrounds. It performs well in crack detection tasks, not only significantly improving the detection accuracy, but also having strong real-time and anti-interference capabilities, providing strong technical support for the health monitoring of hydraulic structures.

[0163] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0164] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

[0165] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting cracks in a hydraulic structure, characterized in that: The method comprises: Acquire an image of a hydraulic structure to be detected, and roughly segment the image to be detected to obtain a roughly segmented image; Performing feature aggregation on the image to be detected and the roughly segmented image to obtain a feature fused image; Performing feature extraction on the feature fusion image using depthwise separable convolution to obtain a feature extraction image; Using a composite attention mechanism to weight the feature extraction image to obtain a feature weighted image; Using an activation function to enhance the features of the feature weighted image to obtain a crack detection result; The expression of feature aggregation is as follows: In the above formula, F agg represents the feature fusion result, represents the number of layers of the input feature map, L represents the total number of layers of the input feature map, F l Indicates The input feature map of the layer, Conv 1×1 represents a 1×1 convolution operation, u l Represents the operation of upsampling input feature maps of different resolutions to the target size; The step of extracting features from the feature fusion image using depthwise separable convolution includes: A deep convolution is performed on each channel of the feature fusion image to obtain a deep convolution result. The formula of the deep convolution is as follows: In the above formula, Represents the output result of deep convolution, i and j Represent the row index and column index on the feature map respectively, c represents the depthwise convolution channel, m and n Respectively represent the width and height of the depth convolution kernel, K Indicates the total amount of offset, Indicates the depth of the convolution kernel at the offset ( m , n ) and in the channel c The weight on The depth convolution result is convolved point by point, and the formula of the point by point convolution is as follows: In the above formula, Represents the point-by-point convolution output result, k represents the point-wise convolution channel, Represents the depth convolution channel c With point-wise convolution channels k The weight between C Indicates the total number of channels; The step of using an activation function to enhance the features of the feature weighted image comprises: The Swish activation function is used to enhance the crack edge details in the feature weighted image to obtain an initial detection image; The Mish activation function is used to converge the crack edge features in the initial detection image.

2. The hydraulic structure crack detection method according to claim 1, characterized in that: The method of weighting the feature extraction image using a composite attention mechanism includes: Using channel attention and spatial attention to weight the feature extraction image in the spatial dimension to obtain an initial weighted feature map; Performing global average pooling on the initial weighted feature map and generating an adaptive weight vector; The channel importance of the initial weighted feature map is weighted according to the adaptive weight vector.

3. The hydraulic structure crack detection method according to claim 2, characterized in that: The step of using channel attention and spatial attention to weight the feature extraction image in a spatial dimension includes: The feature extraction image is weighted using channel attention, and the expression for weighting the channel attention is as follows: In the above formula, M c represents the channel attention weight, Represents the Sigmoid function, Re LU represents a nonlinear activation function, AVGPool ( F ) represents the global maximum pooling of the input feature map. MaxPool ( F ) represents global average pooling of the input feature map. W c1 and W c2 Respectively represent the weight matrix; The feature extraction image is weighted using spatial attention, and the expression for weighting the spatial attention is as follows: In the above formula, M s represents the spatial attention weight, Conv 7×7 Represents a 7×7 convolution operation.

4. The method for detecting cracks in hydraulic structures according to claim 3, characterized in that: The step of weighting the channel importance of the initial weighted feature map according to the adaptive weight vector comprises: Channel attention is used to weight the importance of the channel. The expression for weighting the channel attention is as follows: In the above formula, M se represents the channel importance weight, Represents the Sigmoid function, Re LU represents a nonlinear activation function, represents the adaptive weight vector, W se1 and W se2 They represent weight matrices respectively.

5. A hydraulic structure crack detection device, characterized in that: The device comprises: An image processing module is used to obtain an image of a hydraulic structure to be detected, and to roughly segment the image to be detected to obtain a roughly segmented image; A feature aggregation module, used for performing feature aggregation on the image to be detected and the coarse segmentation image to obtain a feature fusion image; A feature extraction module, used to extract features from the feature fusion image using depthwise separable convolution to obtain a feature extraction image; A feature weighting module, used for weighting the feature extraction image by using a composite attention mechanism to obtain a feature weighted image; A crack detection module, used for enhancing the features of the feature weighted image by using an activation function to obtain a crack detection result; The expression of feature aggregation is as follows: In the above formula, F agg represents the feature fusion result, represents the number of layers of the input feature map, L represents the total number of layers of the input feature map, F l Indicates The input feature map of the layer, Conv 1×1 represents a 1×1 convolution operation, u l Represents the operation of upsampling input feature maps of different resolutions to the target size; The step of extracting features from the feature fusion image using depthwise separable convolution includes: A deep convolution is performed on each channel of the feature fusion image to obtain a deep convolution result. The formula of the deep convolution is as follows: In the above formula, Represents the output result of deep convolution, i and j Represent the row index and column index on the feature map respectively, c represents the depth convolution channel, m and n Respectively represent the width and height of the depth convolution kernel, K Indicates the total amount of offset, Indicates the depth of the convolution kernel at the offset ( m , n ) and in the channel c The weight on The depth convolution result is convolved point by point, and the formula of the point by point convolution is as follows: In the above formula, Represents the point-by-point convolution output result, k represents the point-wise convolution channel, Represents the depth convolution channel c With point-wise convolution channels k The weight between C Indicates the total number of channels; The step of using an activation function to enhance the features of the feature weighted image comprises: The Swish activation function is used to enhance the crack edge details in the feature weighted image to obtain an initial detection image; The Mish activation function is used to converge the crack edge features in the initial detection image.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: The processor executes the computer program to implement the hydraulic structure crack detection method according to any one of claims 1 to 4.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the hydraulic structure crack detection method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Pavement crack detection method and device and electronic equipment

    CN118761964A

  • Bridge crack detection method based on multi-scale feature fusion and multi-layer attention

    CN118941542A