An underwater dam defect detection method, system and medium
By improving the encoder-decoder model and combining bidirectional collaborative attention and dual-domain feature aggregation, the problems of low efficiency and insufficient accuracy of traditional underwater dam detection methods are solved, achieving efficient and accurate underwater dam defect detection that is adaptable to complex underwater environments.
Patent Information
- Application Number
- CN202510367899.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-03-26
AI Technical Summary
Traditional underwater dam defect detection methods are inefficient and lack accuracy, and are difficult to adapt to problems such as noise, ambiguity and uneven lighting in complex underwater environments, thus failing to meet the needs of modern engineering structure health monitoring.
An encoder-decoder model is adopted, replacing the CASS module in the encoder with a bidirectional collaborative attention module, and replacing the synthesis module with a dual-domain feature aggregation module. By combining Zigzag feature scanning and dual-domain feature aggregation, the model is optimized through adaptive global feature selection and a weighted combination of binary cross-entropy loss and Dice loss function.
It significantly improves the accuracy and efficiency of defect detection, effectively addresses noise, ambiguity, and uneven lighting in underwater environments, enhances defect detection capabilities in complex underwater environments, and provides an efficient and accurate detection solution.
Smart Images

Figure CN120163814B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of underwater dam defect detection, in particular to an underwater dam defect detection method, system and medium. BACKGROUND
[0002] In modern society, the monitoring and monitoring of critical infrastructure is crucial, and concrete dams, as the core facilities for flood control, water supply, irrigation, navigation and hydropower generation, play an irreplaceable role. However, due to the influence of external static load and dynamic load such as long-term hydrostatic pressure, temperature load, and the combined effect of internal material aging, concrete dams are prone to defects, which significantly increases the risk of structural failure. These defects not only damage the structural integrity of the dam and change the stress distribution, but also may extend to the interior under water pressure, causing catastrophic consequences.
[0003] Therefore, timely and accurate detection of defects and taking protective measures are crucial for ensuring dam safety. Traditional underwater dam defect detection mainly relies on manual visual inspection, i.e. visual assessment by divers. However, this method has low efficiency, large subjective errors, is significantly affected by environmental conditions, and has high personal safety risks, making it difficult to meet the needs of modern engineering structure health monitoring. In recent years, computer vision and deep learning-based technologies have been gradually introduced into the defect detection field, but due to the complexity of underwater environments (such as uneven lighting, noise interference, image blur, etc.), traditional algorithms perform poorly in practical applications. Therefore, developing an efficient, accurate and adaptive dam defect automatic detection method for complex underwater environments has become a key problem in the field of engineering structure health monitoring. SUMMARY
[0004] The purpose of the present application is to overcome the deficiencies in the prior art and provide an underwater dam defect detection method, system and medium that can effectively deal with noise, blur and uneven lighting in underwater environments, providing reliable technical support for the safety monitoring and maintenance of underwater dams.
[0005] To achieve the above-mentioned purpose, the present application is implemented by using the following technical solutions:
[0006] On the one hand, the present application provides an underwater dam defect detection method, comprising:
[0007] obtaining an original underwater dam image;
[0008] inputting the original underwater dam image into a pre-constructed defect detection model to output a dam defect detection result;
[0009] wherein the construction of the defect detection model comprises:
[0010] Obtaining an encoder-decoder model;
[0011] The output layer of the CASS module in the encoder is replaced by a bidirectional collaborative attention module, and the synthesis module is replaced by a dual-domain feature aggregation module, to obtain a constructed defect detection model.
[0012] Optionally, the processing step of the defect detection model comprises:
[0013] In the encoder, Zigzag feature scanning is performed on the original underwater dam image to obtain scanning features, and bidirectional collaborative attention module is used to extract dual-branch features from the scanning features to obtain an output feature map;
[0014] In the dual-domain feature aggregation module, dual-domain feature aggregation is performed on the output feature map to obtain aggregated features;
[0015] In the encoder, the aggregated features are decoded and reconstructed to obtain a dam defect detection result. Optionally, Zigzag feature scanning is performed on the original underwater dam image to obtain scanning features, comprising:
[0016] Continuous row feature scanning is performed on the original underwater dam image to obtain horizontal scanning features;
[0017] Continuous column feature scanning is performed on the original underwater dam image to obtain vertical scanning features;
[0018] The original underwater dam image is divided into multiple windows, and feature scanning is performed on each window to obtain local scanning features.
[0019] Optionally, it further comprises:
[0020] When the original underwater dam image is scanned to the end of a row or a column, the direction is reversed to continue feature scanning on the position adjacent to the current scanning position according to a preset scanning position mapping table, until all positions are scanned.
[0021] Optionally, the dual-branch feature extraction is performed on the scanning features to obtain an output feature map, comprising:
[0022] ;
[0023] ;
[0024] ;
[0025] ;
[0026] ;
[0027] ;
[0028] ;
[0029] ;
[0030] wherein, represents the feature obtained after grouping the input feature map; represents the feature obtained after height-wise average pooling; represents the feature obtained after width-wise average pooling; represents the feature obtained after further processing; represents the feature obtained after further processing; represents the feature obtained after further processing; represents the feature obtained after further processing; represents the feature obtained after further processing; represents the feature obtained after further processing; represents the feature obtained after further processing; , , element-wise multiplication; represents the result obtained after convolution on ; represents the generated attention weight matrix; represents the output feature map; represents the input feature map; represents the input feature map; represents the number of groups; represents the group dimension transformation operation; represents the result of height-wise average pooling; represents the result of width-wise average pooling; represents the th sub-segment of ; represents the number of sub-segments; represents the one-dimensional convolution operation using a convolution kernel size of ; represents concatenating multiple feature dimensions; represents different convolution kernel sizes; represents GroupNorm processing; represents the Sigmoid activation function; represents the th sub-segment of ; represents convolution; represents the softmax activation function; represents the image height and width; represents element-wise multiplication; represents matrix multiplication.
[0031] Optionally, the output feature map is subjected to double-domain feature aggregation to obtain aggregated features, comprising:
[0032] ;
[0033] ;
[0034] ;
[0035] ;
[0036] ;
[0037] ;
[0038] ;
[0039] ;
[0040] ;
[0041] ;
[0042] ;
[0043] ;
[0044] wherein, represents features fused after grouped convolution and normalization; represents features after double average pooling; represents detail features captured by the difference branch; represents features extracted by the product branch; represents features after merging; represents features after spatial domain aggregation; represents fused features after concatenation and depth separable convolution of low-frequency components; represents fused features after concatenation and point convolution of high-frequency components; represents features after wavelet reconstruction and residual connection; represents encoder input features; represents decoder input features; represents element-wise addition; represents 3x3 grouped convolution; represents batch normalization; represents an activation function; represents average pooling; represents convolution operation; concatenates multiple feature dimensions; represents point-wise convolution; represents low-frequency component of encoder feature; represents horizontal high-frequency component of encoder feature; represents vertical high-frequency component of encoder feature; represents diagonal high-frequency component of encoder feature; represents low-frequency component of decoder feature; represents horizontal high-frequency component of decoder feature; represents vertical high-frequency component of decoder feature; represents diagonal high-frequency component of decoder feature; represents discrete wavelet transform; represents deep convolution; represents inverse discrete wavelet transform; represents output feature map, represents adaptive global feature selection.
[0045] Optionally, the training of the defect detection model comprises:
[0046] The defect detection model is trained by using a weighted sum of a binary cross-entropy loss function and a Dice loss function as a combined loss function, to obtain the trained defect detection model.
[0047] Optionally, the formula of the combined loss function is:
[0048] ;
[0049] ;
[0050] ;
[0051] wherein, represents a binary cross-entropy loss function; represents a Dice loss function; represents a combined loss function; represents a result predicted by a segmentation network; represents a true value corresponding to an input image.
[0052] In a second aspect, the present application provides an underwater dam defect detection system, comprising:
[0053] an image acquisition module, configured to acquire an original underwater dam image;
[0054] an imaging detection module, configured to input the original underwater dam image into a pre-constructed defect detection model, and output a dam defect detection result;
[0055] The model processing module, wherein the construction of the defect detection model includes:
[0056] Obtain the encoder-decoder model;
[0057] The output layer of the CASS module in the encoder is replaced with a bidirectional collaborative attention module, and the synthesis module is replaced with a dual-domain feature aggregation module to obtain the constructed defect detection model.
[0058] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the underwater dam defect detection method described in the first aspect.
[0059] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:
[0060] This invention improves the accuracy and efficiency of defect detection through an innovative model, effectively addressing issues such as noise, ambiguity, and uneven lighting in underwater environments. It provides an efficient and accurate solution for underwater dam defect detection, possessing significant engineering application value. The bidirectional collaborative attention module extracts local and global features, fusing horizontal, vertical, and local block scans using the Zigzag scanning method to comprehensively extract defect features, adapting to complex underwater environments. The dual-domain feature aggregation module optimizes features through spatial and frequency domains, combined with adaptive global feature selection, enhancing defect edge sensitivity and model robustness. Attached Figure Description
[0061] Figure 1 The diagram shown is a flowchart of one embodiment of the underwater dam defect detection method of the present invention;
[0062] Figure 2 The diagram shown is a structural schematic of the bidirectional collaborative attention module in one embodiment of the present invention;
[0063] Figure 3 The diagram shown is a schematic representation of the Zigzag scanning method in one embodiment of the present invention.
[0064] Figure 4 The illustration shows a 2D selective scanning method in one embodiment of the present invention;
[0065] Figure 5 The diagram shown is a structural schematic of a dual-domain feature aggregation module in one embodiment of the present invention;
[0066] Figure 6 The diagram shown illustrates the principle of Haar wavelet transform decomposition in one embodiment of the present invention.
[0067] Figure 7 The diagram shown is a structural schematic of the adaptive global feature selection module in one embodiment of the present invention. Detailed Implementation
[0068] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations thereof. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0069] The term "and / or" simply describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0070] Example 1
[0071] like Figure 1 As shown in the figure, this embodiment introduces an underwater dam defect detection method, which can solve the problems of low efficiency, insufficient accuracy and poor adaptability to complex environments of traditional underwater dam defect detection methods. It can effectively deal with problems such as noise, ambiguity and uneven lighting in the underwater environment, and provide reliable technical support for the safety monitoring and maintenance of underwater dams.
[0072] The method includes the following steps:
[0073] Step 1: Obtain the original underwater dam image, specifically:
[0074] A real-world dataset specifically designed for underwater dam defect detection was constructed, containing 250 images taken directly from the underwater dam environment. Each image was selected frame by frame, and defect instances were annotated using annotation tools to ensure data accuracy.
[0075] Data augmentation was performed on the original underwater dam images, including operations such as rotation, cropping, and horizontal flipping, to simulate the appearance of defects at different angles and scales, increasing data diversity.
[0076] The dataset ultimately contains 1216 images, each with a uniform resolution of 448×448 pixels. This self-built dataset provides important data support for underwater defect detection tasks.
[0077] Step 2: Input the original underwater dam defect image into the pre-built defect detection model, and output the dam defect detection results, specifically:
[0078] The defect detection model is an improvement on the encoder-decoder model. The output layer of the CASS module in the encoder is replaced with a bidirectional collaborative attention module, and the synthesis module is replaced with a dual-domain feature aggregation module, resulting in the constructed defect detection model.
[0079] The encoder progressively constructs hierarchical feature representations through a multi-level network, with each stage containing downsampling layers and multiple cascaded attention state space modules. Encoder and decoder features are fused using a dual-domain feature aggregation module. Finally, the decoder optimizes encoder features using a multi-scale convolution module and performs point convolution and upsampling on the outputs of each decoder branch, adding them element-wise to generate the final detection result. This multi-branch fusion strategy significantly improves multi-scale feature representation capabilities, enhancing defect detection accuracy and consistency.
[0080] First, in the encoder's state-space model (SSM), to address the limitations of traditional scanning methods in defect feature extraction, this embodiment proposes an enhanced multi-mode Zigzag feature scanning strategy, such as... Figure 3 As shown, this strategy combines three feature scanning methods—horizontal, vertical, and segmented local—to comprehensively extract defect features from different angles, resulting in horizontal scanning features, vertical scanning features, and local scanning features:
[0081] like Figure 4 As shown, horizontal feature scanning captures horizontally extending defect features by performing continuous row feature scanning on the original underwater dam image, reducing information breaks caused by skipping rows; vertical feature scanning targets vertical defects by performing continuous column feature scanning on the original underwater dam image, avoiding the weakening problem of vertical correlation in traditional horizontal scanning; block-based local feature scanning divides the original underwater dam image into multiple local windows, performing high-density scanning within each window to enhance sensitivity to minute defects and complex topological structures;
[0082] During the scanning process, when the original underwater dam image is scanned to the end of a row or column, the scanning direction is reversed according to the preset scanning position mapping table to continue feature scanning of the position adjacent to the current scanning position until all positions have been scanned. This strategy significantly improves the ability to capture defect features while ensuring scanning efficiency, and is especially suitable for defect detection tasks in complex underwater environments.
[0083] Secondly, such as Figure 2As shown, in the encoder's Dual-Branch Collaborative Attention (DBCA) module, the input feature maps are first grouped to reduce the number of parameters and avoid excessive coupling between channels. Secondly, bidirectional pooling is used to capture feature responses in the horizontal and vertical directions, enhancing the linear structural characteristics of defects. Then, the pooling results are divided into multiple segments, and multi-scale features are captured through one-dimensional convolution operations with different receptive fields. The feature contribution is adaptively adjusted, thereby significantly improving the ability to extract and represent defect features in complex underwater environments.
[0084] ;
[0085] ;
[0086] ;
[0087] ;
[0088] The DBCA module is a dual-branch feature extraction module. It performs dual-branch feature extraction on the scanned features to obtain an output feature map. In the main branch, weighted features are fused, and directional attention and global information are used to enhance structural consistency. The auxiliary branch extracts detailed features through local convolutions to correct local response biases. Finally, collaborative attention fusion is achieved through cross-branch matrix multiplication to generate the final attention map, which is then multiplied with the grouped feature maps to obtain the output. This module, through a dual-branch complementary mechanism, balances global semantics and local details, significantly improving defect detection performance.
[0089] ;
[0090] ;
[0091] ;
[0092] ;
[0093] ;
[0094] ;
[0095] ;
[0096] ;
[0097] in, This indicates that the features are obtained after grouping the input feature map; Indicates to Features obtained by average pooling over height; Indicates to Features obtained by average pooling across the width; Indicates to Further processing yields the desired features; Indicates to Further processing yields the desired features; Indicates to , , Features obtained by element-wise multiplication; Indicates to , , Element-wise multiplication yields a feature with height i and width j; express Feature weights normalized using Softmax; express Feature weights normalized using Softmax; express Features obtained through global average pooling; express Features obtained through global average pooling; Indicates to conduct The result obtained after convolution; Indicates to conduct After convolution, we get a result with height i and width j; This represents the generated attention weight matrix; This represents the output feature map; Indicates the input feature map; Indicates the number of groups; This indicates a grouping dimension transformation operation; This represents the average pooling result in the height dimension; This represents the average pooling result along the width dimension; express The Each segment; Indicates the number of sub-fields ; This indicates that a convolution kernel size of 1 is used. One-dimensional convolution operation; This indicates that multiple feature dimensions are concatenated; Indicates different convolution kernel sizes; Indicates GroupNorm processing; This represents the Sigmoid activation function; express The Each segment; express convolution; This represents the softmax activation function; Indicates the height and width of the image; This indicates element-wise multiplication; This represents matrix multiplication.
[0098] Then, as Figure 5 As shown, in the dual-domain feature aggregation module, in order to overcome the limitations of traditional skip connections in underwater defect detection tasks, this embodiment proposes a dual-domain feature aggregation module. This module optimizes features from the spatial and frequency dimensions respectively through a spatial domain feature aggregation module and a frequency domain feature aggregation module. Then, through the Adaptive Global Feature Selection (AGFS) module, semantic alignment of spatial and frequency domain features is achieved, thereby optimizing the foreground feature extraction accuracy.
[0099] The spatial domain feature aggregation module adopts a dual-branch design. The difference branch captures subtle differences in edge details, while the product branch strengthens macroscopic region feature modeling, significantly improving defect edge sensitivity and structural continuity recognition capabilities.
[0100] ;
[0101] ;
[0102] ;
[0103] ;
[0104] ;
[0105] ;
[0106] The frequency domain feature aggregation module uses wavelet transform to decompose features into high- and low-frequency components, processing low-frequency and high-frequency information separately, such as... Figure 6 The diagram shown illustrates the principle of Haar wavelet transform decomposition, which effectively suppresses noise interference and enhances adaptability to irregular defects.
[0107] ;
[0108] ;
[0109] ;
[0110] ;
[0111] ;
[0112] like Figure 7 As shown, the AGFS module achieves semantic alignment between spatial domain features and frequency domain features, optimizing the accuracy of foreground feature extraction, i.e.:
[0113] ;
[0114] in, This represents the features fused after grouped convolution and normalization; This represents the features after double average pooling; This indicates the detailed features captured by the difference branch; This represents the features extracted by the product branch; Indicates the characteristics after merging; This represents the features after spatial domain aggregation; This represents the fused features of low-frequency components after splicing and depthwise separable convolution; This represents the fusion feature of high-frequency components after splicing and point convolution; This represents the features after wavelet reconstruction and residual connection; Indicates encoder input features; Indicates the input features of the decoder; This indicates element-wise addition; This represents a 3×3 grouped convolution; Indicates batch normalization; Indicates the activation function; Indicates average pooling; Indicates the convolution operation; This indicates that multiple feature dimensions are concatenated; This represents pointwise convolution; This represents the low-frequency component of the encoder characteristics; Represents the high-frequency components of the encoder feature level; This represents the vertical high-frequency component of the encoder feature; This represents the diagonal high-frequency components characteristic of the encoder; This indicates the low-frequency component characteristic of the decoder; This represents the high-frequency components of the decoder's features. This indicates the vertical high-frequency component of the decoder. This represents the diagonal high-frequency components characteristic of the decoder; Represents the discrete wavelet transform; Represents depthwise convolution; Represents the discrete inverse wavelet transform; This represents the output feature map. This indicates adaptive global feature selection.
[0115] The DBCA module takes into account both global and local features, significantly improving the model's ability to capture defect features in complex underwater environments and its robustness. This innovative design provides a powerful and efficient solution for underwater defect detection tasks.
[0116] Finally, the aggregated features are decoded and reconstructed using a decoder to obtain the dam defect detection results.
[0117] The training of the defect detection model includes:
[0118] We employ a weighted sum of Binary Cross-Entropy Loss (BCE Loss) and Dice Loss as the model's loss function. Since defect detection essentially classifies pixels into defective and non-defective pixels, making it a binary classification problem, we use BCE Loss in the loss function. BCE Loss provides fine-grained optimization signals by comparing the predicted results with the ground truth labels pixel-by-pixel. Furthermore, in underwater dam defect detection, severe class imbalance exists, which may cause the model to overly favor non-defective pixels. Therefore, we introduce Dice Loss, which effectively alleviates the positive and negative class imbalance problem by directly optimizing the region similarity between the predicted results and the ground truth labels.
[0119] To leverage the advantages of both loss functions, we use a weighted combination of them. This combined loss function effectively addresses class imbalance while simultaneously optimizing the accuracy and boundary consistency of detection results. It improves pixel-level classification accuracy and exhibits strong robustness in handling ambiguous boundary issues. The formula for calculating the combined loss function is as follows:
[0120] ;
[0121] ;
[0122] ;
[0123] in, Represents the binary cross-entropy loss function; Represents the Dice loss function; Represents the combined loss function; This represents the result predicted by the segmentation network; This represents the ground truth value corresponding to the input image.
[0124] In this embodiment, the batch size is 8, the total training cycle is 300, the Adam algorithm is used to optimize the model parameters, and a warm-up technique is used to ensure that the learning rate gradually increases to the set value to avoid early overfitting.
[0125] The defect detection method in this embodiment performs excellently in underwater dam image defect detection: the model IoU reaches 75.74% and the Dice reaches 86.19%, which is better than methods such as Unet, DeepLabv3+, FADC, Mask2former, and PEM. It shows unique detail and adaptability in small-sized defect detection and effectively identifies defects of different shapes and sizes in complex underwater environments.
[0126] Example 2
[0127] Based on Example 1, this example introduces an underwater dam defect detection system, including:
[0128] The image acquisition module acquires raw underwater images of the dam.
[0129] The imaging detection module inputs the original underwater dam image into a pre-built defect detection model and outputs the dam defect detection results.
[0130] The model building module, wherein the construction of the defect detection model includes:
[0131] Obtain the encoder-decoder model;
[0132] The output layer of the CASS module in the encoder is replaced with a bidirectional collaborative attention module, and the synthesis module is replaced with a dual-domain feature aggregation module to obtain the constructed defect detection model.
[0133] The specific functions of each module described above are explained in the relevant content of Embodiment 1 or 2, and will not be repeated here.
[0134] Example 3
[0135] This embodiment introduces a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the underwater dam defect detection method described in Embodiment 1.
[0136] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0137] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0138] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0139] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0140] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.
Claims
1. A method for detecting defects in underwater dams, characterized in that, include: Obtain raw underwater images of the dam; The original underwater dam image is input into a pre-built defect detection model, and the dam defect detection results are output. The construction of the defect detection model includes: Obtain the encoder-decoder model; The encoder includes a CASS module, which is based on the VSS module and adds an LN layer, a two-layer convolutional layer, and a bidirectional collaborative attention module connected in sequence after the VSS module. The skip connections in the encoder-decoder model are replaced with a dual-domain feature aggregation module to obtain the constructed defect detection model. In the encoder, the original underwater dam image is scanned using Zigzag features to obtain scan features. The bidirectional collaborative attention module is a two-branch feature extraction module that extracts features from the scan features in two branches to obtain an output feature map. In the bidirectional collaborative attention module, the scan features are grouped to obtain a grouped feature map. In the main branch, weighted features are fused, and bidirectional pooling is used to capture horizontal and vertical feature responses respectively, enhancing the linear structural characteristics of the defects. The pooling result is divided into multiple segments, and one-dimensional convolution operations with different receptive fields are used to capture multi-scale features. Directional attention and global information are used to enhance structural consistency. The auxiliary branch extracts detailed features through local convolution to correct local response biases. Finally, collaborative attention fusion is achieved through cross-branch matrix multiplication to generate the final attention map, which is then multiplied with the grouped feature map to obtain the output feature map. In the dual-domain feature aggregation module, the output feature map is subjected to dual-domain feature aggregation to obtain aggregated features. The dual-domain feature aggregation module optimizes the features from the spatial and frequency dimensions respectively through the spatial domain feature aggregation module and the frequency domain feature aggregation module. Then, the semantic alignment of the spatial domain and frequency domain features is achieved through the adaptive global feature selection module.
2. The underwater dam defect detection method according to claim 1, characterized in that, The original underwater dam image was subjected to Zigzag feature scanning to obtain scanning features, including: The original underwater dam image was subjected to continuous row feature scanning to obtain lateral scan features; The original underwater dam image is subjected to continuous column feature scanning to obtain longitudinal scan features; The original underwater dam image is divided into multiple windows, and feature scanning is performed on each window to obtain local scan features.
3. The underwater dam defect detection method according to claim 2, characterized in that, Also includes: When the original underwater dam image is scanned to the end of a row or column, according to the preset scan position mapping table, the direction is reversed to continue feature scanning of the position adjacent to the current scan position until feature scanning of all positions is completed.
4. The underwater dam defect detection method according to claim 1, characterized in that, The scanned features are subjected to bi-branch feature extraction to obtain an output feature map, including: ; ; ; ; ; ; ; ; in, This indicates that the features are obtained after grouping the input feature map; Indicates to Features obtained by average pooling over height; Indicates to Features obtained by average pooling across the width; Indicates to Further processing yields the desired features; Indicates to Further processing yields the desired features; Indicates to , , Features obtained by element-wise multiplication; Indicates to conduct The result obtained after convolution; This represents the generated attention weight matrix; This represents the output feature map; Indicates the input feature map; Indicates the number of groups; This indicates a grouping dimension transformation operation; This represents the average pooling result in the height dimension; This represents the average pooling result along the width dimension; express The Each segment; Indicates the number of sub-fields; This indicates that a convolution kernel size of 1 is used. One-dimensional convolution operation; This indicates that multiple feature dimensions are concatenated; Indicates different convolution kernel sizes; Indicates GroupNorm processing; This represents the Sigmoid activation function; express The Each segment; express convolution; This represents the softmax activation function; Indicates the height and width of the image; This indicates element-wise multiplication; This represents matrix multiplication.
5. The underwater dam defect detection method according to claim 1, characterized in that, The output feature map is subjected to dual-domain feature aggregation to obtain aggregated features, including: ; ; ; ; ; ; ; ; ; ; ; ; in, This represents the features fused after grouped convolution and normalization; This represents the features after double average pooling; This indicates the detailed features captured by the difference branch; This represents the features extracted by the product branch; Indicates the characteristics after merging; This represents the features after spatial domain aggregation; This represents the fused features of low-frequency components after splicing and depthwise separable convolution; This represents the fusion feature of high-frequency components after splicing and point convolution; This represents the features after wavelet reconstruction and residual connection; Indicates encoder input features; Indicates the input features of the decoder; This indicates element-wise addition; This represents a 3×3 grouped convolution; Indicates batch normalization; Indicates the activation function; Indicates average pooling; Indicates the convolution operation; This indicates that multiple feature dimensions are concatenated; This represents pointwise convolution; This represents the low-frequency component of the encoder characteristics; Represents the high-frequency components of the encoder feature level; This represents the vertical high-frequency component of the encoder feature; This represents the diagonal high-frequency components characteristic of the encoder; This indicates the low-frequency component characteristic of the decoder; This represents the high-frequency components of the decoder's features. This indicates the vertical high-frequency component of the decoder. This represents the diagonal high-frequency components characteristic of the decoder; Represents the discrete wavelet transform; Represents depthwise convolution; Represents the discrete inverse wavelet transform; This represents the output feature map. This indicates adaptive global feature selection.
6. The underwater dam defect detection method according to claim 1, characterized in that, The training of the defect detection model includes: The defect detection model is trained by using a weighted sum of the binary cross-entropy loss function and the Dice loss function as a combined loss function, resulting in a well-trained defect detection model.
7. The underwater dam defect detection method according to claim 6, characterized in that, The formula for calculating the combined loss function is as follows: ; ; ; in, Represents the binary cross-entropy loss function; Represents the Dice loss function; Represents the combined loss function; This represents the result predicted by the segmentation network; This represents the true value corresponding to the input image.
8. An underwater dam defect detection system, characterized in that, include: The image acquisition module acquires raw underwater images of the dam. The imaging detection module inputs the original underwater dam image into a pre-built defect detection model and outputs the dam defect detection results. The model building module, wherein the construction of the defect detection model includes: Obtain the encoder-decoder model; The encoder includes a CASS module, which is based on the VSS module and adds an LN layer, a two-layer convolutional layer, and a bidirectional collaborative attention module connected in sequence after the VSS module. The skip connections in the encoder-decoder model are replaced with a dual-domain feature aggregation module to obtain the constructed defect detection model. In the encoder, the original underwater dam image is scanned using Zigzag features to obtain scan features. The bidirectional collaborative attention module is a two-branch feature extraction module that extracts features from the scan features in two branches to obtain an output feature map. In the bidirectional collaborative attention module, the scan features are grouped to obtain a grouped feature map. In the main branch, weighted features are fused, and bidirectional pooling is used to capture horizontal and vertical feature responses respectively, enhancing the linear structural characteristics of the defects. The pooling result is divided into multiple segments, and one-dimensional convolution operations with different receptive fields are used to capture multi-scale features. Directional attention and global information are used to enhance structural consistency. The auxiliary branch extracts detailed features through local convolution to correct local response biases. Finally, collaborative attention fusion is achieved through cross-branch matrix multiplication to generate the final attention map, which is then multiplied with the grouped feature map to obtain the output feature map. In the dual-domain feature aggregation module, the output feature map is subjected to dual-domain feature aggregation to obtain aggregated features. The dual-domain feature aggregation module optimizes the features from the spatial and frequency dimensions respectively through the spatial domain feature aggregation module and the frequency domain feature aggregation module. Then, the semantic alignment of the spatial domain and frequency domain features is achieved through the adaptive global feature selection module.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the underwater dam defect detection method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Video code rate control method based on multi-scale attention content perception
CN118354080A
Building outer wall crack detection method and system, medium and program product
CN119006454A