A method for detecting and identifying defects of a dam

CN122820583APending Publication Date: 2026-09-25HOHAI UNIV SUZHOU RES INST +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610923447.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-25
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]本发明的目的是提供一种大坝缺陷检测与识别方法,通过对待检测的大坝标准化图像进行逐层下采样与上采样融合多尺度特征,对三维张量形式的空间域图像特征进行频域低秩分解与张量奇异谱分解,并采用锚框对融合特征张量进行缺陷分类与边界框回归,得到缺陷检测与识别结果,以解决复杂水下环境下大坝缺陷特征易被噪声掩盖、定位不准的问题

Benefits of technology

[0066]1、本发明通过获取大坝渗漏红外图像和水下大坝结构图像并进行预处理,对大坝标准化图像逐层下采样提取浅层纹理与深层语义后执行上采样融合多尺度特征,输出三维张量形式的空间域图像特征,并对所述空间域图像特征进行频域低秩分解与张量奇异谱分解,得到融合特征张量,利用所述融合特征张量进行缺陷分类与边界框回归,实现红外与水下两类模态图像的统一特征建模与跨域特征融合,有效抑制红外杂波与水体散射噪声对缺陷特征的掩盖,提升了复杂水下环境下大坝缺陷的检测精度,解决了复杂水下环境下大坝缺陷特征易被噪声掩盖、定位不准的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820583A_ABST
    Figure CN122820583A_ABST
Patent Text Reader

Abstract

The application discloses a dam defect detection and identification method, and belongs to the technical field of water conservancy safety monitoring. The method comprises the following steps: obtaining a dam multimodal original image to be detected and preprocessing the dam multimodal original image to obtain a dam standardized image, processing based on a dam defect detection model, and outputting a defect detection and identification result; and a data processing method of the dam defect detection model, which comprises the following steps: performing layer-by-layer downsampling on the dam standardized image to extract shallow texture and deep semantics, performing upsampling on output features, fusing multi-scale features to restore spatial resolution, and outputting spatial domain image features in the form of a three-dimensional tensor; performing frequency domain low-rank decomposition and tensor singular spectrum decomposition on the spatial domain image features to obtain a fusion feature tensor; and using an anchor box to perform defect classification and boundary box regression on the fusion feature tensor to obtain the defect detection and identification result. The application solves the problems that dam defect features are easily covered by noise and are not accurately positioned in a complex underwater environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for detecting and identifying defects in dams, belonging to the field of water conservancy safety monitoring technology. Background Technology

[0002] In recent years, dams, as core infrastructure of water conservancy and hydropower projects, have seen their structural safety directly impacting the safety of life and property downstream and regional economic development. Long-term influence from hydrological environment, geological conditions, material aging, and external loads makes dams prone to defects such as leakage, cracks, concrete spalling, and underwater structural damage. If these defects are not detected and addressed in a timely manner, they can continue to develop and lead to catastrophic accidents such as dam instability and collapse. Currently, dam defect detection mainly includes two core methods: infrared leakage detection and underwater structural detection. Infrared detection uses thermal imaging equipment to collect images of the dam surface temperature distribution, while underwater structural detection uses high-definition underwater cameras to collect images of the dam's underwater structure. In recent years, large-scale artificial intelligence models have demonstrated powerful general feature learning capabilities in the field of computer vision, and some research has begun to attempt to apply them to the field of water conservancy project safety monitoring. However, due to the large number of parameters in the large models and the limited computing power of engineering detection terminals, their engineering deployment still faces many challenges. Meanwhile, existing patents related to dam defect detection mostly focus on improving a single lightweight convolutional neural network, extracting features only in the spatial domain or a single frequency domain, without achieving unified feature modeling and cross-domain feature fusion for multimodal detection images.

[0003] However, existing dam defect detection technologies still have the following shortcomings: First, infrared images of dams are affected by shooting distance and angle, resulting in low resolution, high noise, and blurred boundaries of seepage areas. Underwater detection images are affected by water scattering, absorption, and uneven illumination, leading to color distortion and image blurring. Traditional image preprocessing methods are insufficient to effectively improve image quality, causing defect features to be easily masked by complex environmental noise. Second, traditional defect detection models only extract features in the spatial domain or a single frequency domain, resulting in problems such as pattern aliasing, pseudo-components, and distorted trend extraction. They struggle to effectively separate defect features from background noise, exhibiting low accuracy in identifying minute defects and defects in complex backgrounds, and have weak generalization ability, making it difficult to adapt to the detection needs of different dam types and different detection environments. Third, the bounding box regression loss used in existing technologies has an optimization gap with the core evaluation metric IoU. In non-overlapping bounding box scenarios, the gradient is zero, failing to effectively optimize the bounding box positions of difficult-to-locate defects such as thin cracks and minor leaks, leading to insufficient defect localization accuracy. Fourth, existing technologies fail to achieve unified feature modeling for multimodal detection images. The feature extraction models for infrared and underwater images are independent, lacking cross-modal and cross-domain feature fusion capabilities, making it difficult to achieve integrated detection of dam defects across the entire domain. Furthermore, existing technologies do not integrate the anti-aliasing and signal decoupling capabilities of tensor singular spectral decomposition, the feature analysis capabilities of frequency domain low-rank decomposition, or the precise localization capabilities of GIoU (Generalized Intersection over Union). Consequently, the detection accuracy and generalization ability are insufficient to meet the practical needs of dam safety monitoring. Therefore, developing a dam defect detection and identification method that can effectively cope with interference from complex underwater environments and achieve unified feature modeling and precise localization of multimodal images has become an urgent technical problem to be solved in the field of water conservancy engineering safety monitoring. Summary of the Invention

[0004] The purpose of this invention is to provide a method for detecting and identifying dam defects. This method involves fusing multi-scale features through layer-by-layer downsampling and upsampling of the standardized image of the dam to be detected. The spatial domain image features in the form of a three-dimensional tensor are then subjected to frequency domain low-rank decomposition and tensor singular spectrum decomposition. Anchor frames are then used to classify defects and perform bounding box regression on the fused feature tensor to obtain defect detection and identification results. This method aims to solve the problem that dam defect features are easily masked by noise and are inaccurately located in complex underwater environments.

[0005] To solve the above-mentioned technical problems, the present invention is implemented using the following technical solution.

[0006] This invention provides a method for detecting and identifying defects in dams, comprising:

[0007] The original multimodal images of the dam to be detected are acquired and preprocessed to obtain a standardized image of the dam. The original multimodal images of the dam include infrared images of dam leakage and underwater images of the dam structure.

[0008] Based on the standardized image of the dam, the image is processed using a pre-trained dam defect detection model to output defect detection and identification results, including the type, location, and confidence level of the dam defect.

[0009] The data processing method for the dam defect detection model includes:

[0010] The standardized image of the dam is downsampled layer by layer to extract shallow texture and deep semantics, and the downsampling results are used as multi-scale features, with the deepest feature being used as the output feature. The output feature is then upsampled and multi-scale features are fused to restore spatial resolution, outputting spatial domain image features in the form of a three-dimensional tensor.

[0011] The spatial domain image features are subjected to frequency domain low-rank decomposition and tensor singular spectrum decomposition to obtain the fused feature tensor:

[0012] Anchor frames are used to classify defects and regress bounding boxes on the fused feature tensor to obtain defect detection and recognition results.

[0013] Furthermore, the network structure of the dam defect detection model includes:

[0014] The input layer is used to receive the standardized image of the dam;

[0015] The backbone network includes a multi-level stacked encoder and decoder. The encoder is used to downsample the dam-normalized image layer by layer to extract shallow texture and deep semantics and use the downsampling results as multi-scale features, with the deepest feature as the output feature. The decoder is used to upsample the output features and fuse the multi-scale features to restore spatial resolution and output spatial domain image features in the form of a three-dimensional tensor.

[0016] The decomposition module is used to perform frequency domain low-rank decomposition and tensor singular spectrum decomposition on the spatial domain image features to obtain a fused feature tensor.

[0017] The defect detection head is used to classify defects and regress bounding boxes on the fused feature tensor using anchor frames to obtain defect detection and recognition results.

[0018] The output layer is used to output the defect detection and identification results.

[0019] Furthermore, the spatial domain image features are subjected to frequency domain low-rank decomposition and tensor singular spectrum decomposition to obtain a fused feature tensor, including:

[0020] The spatial domain image features are subjected to Fourier basis expansion and transformed to the frequency dimension to obtain frequency domain features;

[0021] The frequency domain features are recombined using third-order Hankel tensor quantization to obtain a third-order tensor;

[0022] Perform tensor singular spectrum decomposition on the third-order tensor to obtain effective features;

[0023] A nonnegative matrix decomposition with cosine similarity regularization constraints and hyperplane projection constraints is used to divide the effective features and the preset low-rank components of the features to obtain the decomposed features.

[0024] The decomposed features are sparsified, and the effects are classified into layers according to the defect morphology, and the fused feature tensor is output.

[0025] Furthermore, the spatial domain image features are subjected to Fourier basis expansion and transformed to the frequency dimension to obtain frequency domain features, including:

[0026] By adding an extended window to the spatial domain image features, the extended spatial domain features are obtained;

[0027] Based on the window function constraint of the boundary of the extended spatial domain features, the windowed spatial domain features are obtained.

[0028] The windowed spatial domain features are subjected to Fourier basis expansion operations to map them to frequency amplitude spectrum and phase sensing basis functions, thus obtaining frequency domain features.

[0029] Furthermore, the frequency domain features are reconstructed using third-order Hankel tensor quantization to obtain a third-order tensor, including:

[0030] Obtain the feature maps of each channel of the frequency domain features to obtain a multi-channel frequency domain feature map;

[0031] Obtain the neighborhood feature blocks of the multi-channel frequency domain feature map at each pixel location to obtain the frequency domain local feature blocks;

[0032] Based on the frequency domain local feature blocks, a third-order Hankel matrix is ​​constructed by stacking them according to the feature channel dimension to obtain the third-order Hankel matrix.

[0033] The feature maps of each channel in the third-order Hankel matrix are fused and recombined into a third-order Hankel frequency domain feature tensor.

[0034] Furthermore, tensor singular spectrum decomposition is performed on the third-order tensor to obtain effective features, including:

[0035] The mean of each dimension is calculated over the entire domain of the third-order tensor, and the mean of the corresponding dimension is subtracted from each element in the third-order tensor to obtain the centered third-order tensor.

[0036] The centered third-order tensor is subjected to modal expansion operations along the spatial dimension, neighborhood dimension, and feature channel dimension, respectively, and is decomposed into three sets of two-dimensional matrices;

[0037] Singular value decomposition is performed on the three sets of two-dimensional matrices respectively to obtain the left singular vector matrix, singular value diagonal matrix and right singular vector matrix corresponding to each set of two-dimensional matrices, wherein the singular values ​​in the singular value diagonal matrix are sorted in descending order of value;

[0038] Based on the noise distribution of the dam standardized image, a singular value threshold is preset, and singular values ​​in each singular value diagonal matrix that are greater than the singular value threshold, as well as the left singular vector matrix and the right singular vector matrix corresponding to the singular value are retained.

[0039] Based on the singular values ​​greater than the singular value threshold in each retained singular value diagonal matrix, and the left and right singular vector matrices corresponding to the singular values, each set of two-dimensional feature matrices is reconstructed in reverse, and each set of two-dimensional feature matrices is reassembled according to the modal structure and dimensional order of the third-order tensor to obtain effective features.

[0040] Furthermore, a nonnegative matrix decomposition with cosine similarity regularization constraints and hyperplane projection constraints is used to divide the effective features and the preset low-rank components of the features, resulting in decomposed features, including:

[0041] Obtain the effective features and the preset low-rank components of the features as the input matrix to be decomposed;

[0042] Apply a nonnegativity constraint to the input matrix to obtain a nonnegativity-constrained input matrix.

[0043] Based on the input matrix after the non-negativity constraint, the target loss function is constructed by weighted combination of F-norm reconstruction loss term and cosine similarity regularization term.

[0044] Based on the target loss function, nonnegative matrix decomposition is performed on the input matrix to obtain the coefficient matrix and the basis matrix;

[0045] Apply hyperplane projection constraints to the coefficient matrix and the basis matrix to obtain a coefficient matrix and a basis matrix that satisfy the hyperplane constraints;

[0046] Based on the coefficient matrix and basis matrix that satisfy the hyperplane constraint, the decomposition features are reconstructed.

[0047] Furthermore, the decomposed features are sparsified, and the effects are stratified and classified according to the defect morphology, outputting a fused feature tensor, including:

[0048] The decomposition features are sparsified to obtain the sparsified decomposition features.

[0049] The amplitude distribution of the sparsified decomposition features in the spatial frequency domain is obtained, and the amplitude distribution is divided into a low-frequency amplitude part and a high-frequency amplitude part according to a preset spatial frequency threshold to obtain low-frequency components and high-frequency components.

[0050] Based on the low-frequency components and the high-frequency components, the sparsified decomposition features are divided into low-frequency defect features and high-frequency defect features to obtain the hierarchical classification features.

[0051] The features after hierarchical classification are fused to output a fused feature tensor.

[0052] Furthermore, anchor frames are used to perform defect classification and bounding box regression on the fused feature tensor to obtain defect detection and recognition results, including:

[0053] Based on the fused feature tensor, the defect location is predicted using anchor frames with preset scale and aspect ratio to obtain an initial predicted bounding box.

[0054] Based on the initial predicted bounding box, the defect category confidence and bounding box coordinate offset of each anchor box are calculated to obtain the classification result and regression result;

[0055] Based on the classification and regression results, the generalized intersection-union loss function is used as the loss function for bounding box regression. The loss value between the predicted bounding box and the true bounding box is calculated to obtain the optimized bounding box regression result.

[0056] Based on the classification results and the optimized bounding box regression results, the defect detection and identification results are output.

[0057] Furthermore, the training method for the dam defect detection model includes:

[0058] The original multimodal images of the dam were acquired and preprocessed to obtain standardized image samples of the dam.

[0059] Based on the standardized image samples of the dam, the standardized image samples are used as input features, the corresponding pre-labeled defect labels are used as output features, and the samples are divided into training set, validation set and test set according to a preset ratio.

[0060] A multi-dimensional hybrid loss function is constructed, which is a weighted sum of binary cross-entropy loss, Dice loss, perceptual loss, generalized cross-union ratio loss, tensor singular spectrum decomposition reconstruction loss and frequency domain low-rank decomposition reconstruction loss. Each loss term is configured with a weight hyperparameter to obtain the total loss function.

[0061] Based on the total loss function and the training set, the dam defect detection model is pre-trained, the weights of the backbone network are frozen, and the decomposition module and defect detection head are iteratively trained to obtain the pre-trained model.

[0062] Based on the pre-trained model, all network layers are unfrozen, and the dam defect detection model is jointly trained end-to-end using a pre-warmed cosine annealing learning rate strategy to obtain a fine-tuned model.

[0063] Based on the fine-tuned trained model and the validation set, K-fold cross-validation is used to select the optimal hyperparameter combination to obtain the cross-validated model.

[0064] Based on the cross-validated model, generalization validation is performed using the test set, and incremental training is conducted on standardized dam image samples with recognition rates below a preset threshold to obtain a trained dam defect detection model.

[0065] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:

[0066] 1. This invention acquires infrared images of dam leakage and underwater dam structure images and performs preprocessing. It then downsamples the standardized dam images layer by layer to extract shallow texture and deep semantics, followed by upsampling and fusion of multi-scale features to output spatial domain image features in the form of a three-dimensional tensor. The spatial domain image features are then subjected to frequency domain low-rank decomposition and tensor singular spectrum decomposition to obtain a fused feature tensor. This fused feature tensor is used for defect classification and bounding box regression, achieving unified feature modeling and cross-domain feature fusion for both infrared and underwater modal images. This effectively suppresses the masking of defect features by infrared clutter and water scattering noise, improving the detection accuracy of dam defects in complex underwater environments and solving the problem of dam defect features being easily masked by noise and inaccurately located in complex underwater environments.

[0067] 2. This invention acquires infrared images of dam leakage and underwater dam structure images and performs preprocessing. The encoder of the backbone network downsamples layer by layer to extract shallow texture and deep semantics, and the decoder upsamples and fuses multi-scale features to output three-dimensional tensor spatial domain image features. The decomposition module performs frequency domain low-rank decomposition and tensor singular spectrum decomposition on the spatial domain image features to obtain fused feature tensors. This achieves unified feature modeling and cross-domain feature fusion of infrared and underwater modal images in the spatial, frequency, and tensor domains. It can effectively suppress the masking of defect features by infrared clutter and water scattering noise, and improve the detection accuracy of dam defects in complex underwater environments.

[0068] 3. This invention adds an extended window to the spatial domain image features and performs Fourier basis expansion to transform them to the frequency dimension. The frequency domain features are then recombined using third-order Hankel tensor quantization to obtain a third-order tensor. The third-order tensor is then subjected to centering, modal expansion, singular value decomposition, filtering of effective singular values ​​based on a preset threshold for noise distribution in the dam standardized image, and inverse reconstruction to obtain effective features after noise removal. This allows for the separation of defect features from the original image that is masked by noise, solving the problem of difficulty in extracting defect features in complex underwater environments.

[0069] 4. This invention obtains decomposed features by performing non-negative matrix decomposition with cosine similarity regularization constraints and hyperplane projection constraints on effective features and preset low-rank components of features. The decomposed features are then sparsified and divided into low-frequency defect features and high-frequency defect features according to spatial frequency thresholds before being fused to obtain a fused feature tensor. This enables hierarchical classification of trend defects and non-stationary fluctuation defects. The defect detection head uses anchor frames with preset scale and aspect ratio to classify defects on the fused feature tensor and uses the generalized intersection-union loss function as the loss function for bounding box regression, thereby improving the bounding box positioning accuracy of slender cracks and micro-leakage. Attached Figure Description

[0070] Figure 1 This is a flowchart illustrating a method for detecting and identifying defects in dams provided in an embodiment of the present invention;

[0071] Figure 2 This is a schematic diagram of the structure of the dam defect detection model provided in an embodiment of the present invention;

[0072] Figure 3 This is a schematic diagram of the data processing process of the decomposition module provided in an embodiment of the present invention;

[0073] Figure 4 This is a schematic diagram of the data processing process of the defect detection head provided in an embodiment of the present invention. Detailed Implementation

[0074] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations thereof. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.

[0075] Example 1

[0076] like Figure 1 As shown in the figure, this embodiment introduces a method for dam defect detection and identification, including:

[0077] Step 1: Obtain the original multimodal image of the dam to be detected and perform preprocessing to obtain the dam standardized image.

[0078] In this embodiment, the multimodal raw images of the dam include infrared images of dam leakage and underwater images of the dam structure. This embodiment obtains standardized dam images by acquiring and preprocessing the infrared images of dam leakage and underwater images of the dam structure. This provides standardized multimodal input data with uniform specifications for subsequent dam defect detection models, enabling joint processing of infrared and underwater images within the same framework.

[0079] Step 2: Based on the standardized dam image, process it using a pre-trained dam defect detection model to output defect detection and identification results.

[0080] In this embodiment, the defect detection and identification results include the type, location, and confidence level of the dam defect. This embodiment achieves end-to-end automated processing of dam defect detection and identification by inputting a standardized image of the dam into a pre-trained dam defect detection model for processing and outputting the type, location, and confidence level of the dam defect.

[0081] Step 2.1: Perform layer-by-layer downsampling on the standardized image of the dam to extract shallow texture and deep semantics, and use the downsampling results as multi-scale features, with the deepest feature as the output feature; perform upsampling on the output feature, and fuse the multi-scale features to restore spatial resolution, outputting spatial domain image features in the form of a three-dimensional tensor.

[0082] This embodiment extracts shallow texture and deep semantics by downsampling the standardized image of the dam layer by layer, and then upsamples the output features and fuses multi-scale features to restore spatial resolution, outputting spatial domain image features in the form of a three-dimensional tensor. While preserving global semantic information, it restores local detail features and provides a feature representation that preserves spatial structure for frequency domain decomposition.

[0083] Step 2.2: Perform frequency domain low-rank decomposition and tensor singular spectrum decomposition on the spatial domain image features to obtain the fused feature tensor.

[0084] This embodiment obtains a fused feature tensor by performing frequency domain low-rank decomposition and tensor singular spectrum decomposition on the spatial domain image features. In the frequency domain, the defect features are separated from infrared clutter and water scattering noise, and the separated effective features are reassembled into a tensor form suitable for subsequent detection, thereby suppressing the masking of defect features by complex underwater environmental noise.

[0085] Step 2.3: Use anchor frames to perform defect classification and bounding box regression on the fused feature tensor to obtain defect detection and recognition results.

[0086] This embodiment uses anchor frames to classify defects and regress bounding boxes in the fused feature tensor to obtain defect detection and recognition results, thereby realizing the classification and location of defect targets in the fused feature tensor.

[0087] Example 2

[0088] Based on the same inventive concept as Embodiment 1, this embodiment describes the implementation steps of a dam defect detection and identification method, including:

[0089] Step 1: Obtain the original multimodal image of the dam to be detected and perform preprocessing to obtain the dam standardized image.

[0090] In this embodiment, the multimodal raw images of the dam include infrared images of dam leakage and underwater images of the dam structure. Specifically, these images were collected from a field measurement of a hydraulic dam, yielding 8000 valid raw images, including 5200 samples with defects and 2800 samples with no defective backgrounds. The infrared images were acquired using thermal imaging equipment to capture the surface temperature distribution of the dam, with a resolution of 640×480. The underwater structure images were acquired using high-definition cameras to capture the underwater structure of the dam, with a resolution of 1280×720. Differentiated preprocessing was applied to the two types of images: the infrared images underwent grayscale normalization, mapping grayscale values ​​to a unified range and a uniform scale of 640×640; the underwater images underwent histogram equalization to enhance contrast and highlight the edge features of structural defects, also with a uniform scale of 640×640. After preprocessing, data enhancement was performed on all images, including random adjustments to brightness and contrast of ±20% and 5%~10% local random occlusion, expanding the total sample size to 24000 images. Finally, the pixel values ​​of the enhanced image are normalized to the range of 0 to 1 to obtain the dam-standardized image.

[0091] Step 2: Based on the standardized dam image, process it using a pre-trained dam defect detection model to output defect detection and identification results.

[0092] In this embodiment, the defect detection and identification results include the type, location, and confidence level of the dam defect. In this embodiment, the network structure of the dam defect detection model is as follows: Figure 2 As shown, it includes an input layer, a backbone network, a decomposition module, a defect detection head, and an output layer. The input layer is used to receive the standardized image of the dam.

[0093] In this embodiment, the network structure of the dam defect detection model is as follows: Figure 2As shown, the system includes an input layer, a backbone network, a decomposition module, a defect detection head, and an output layer. The input layer receives the standardized image of the dam. The backbone network includes a multi-level stacked encoder and decoder. The encoder performs layer-by-layer downsampling on the standardized dam image to extract shallow texture and deep semantics, using the downsampling results as multi-scale features and the deepest feature as the output feature. The decoder performs upsampling on the output features and fuses the multi-scale features to restore spatial resolution, outputting spatial domain image features in the form of a three-dimensional tensor. The decomposition module performs frequency domain low-rank decomposition and tensor singular spectrum decomposition on the spatial domain image features to obtain a fused feature tensor. The defect detection head uses anchor boxes to perform defect classification and bounding box regression on the fused feature tensor to obtain defect detection and recognition results. The output layer outputs the defect detection and recognition results.

[0094] Specifically, in the backbone network, the encoder employs a 5-level stacked downsampling structure. The input is a standardized image of the dam, with each level having a downsampling factor of 2, extracting features step by step. Levels 1 and 2 downsampling focus on shallow texture features, including crack edges, seepage textures, and underwater structure outlines; levels 3 to 5 downsampling focus on deep semantic features, including the overall morphology of defects and regional distribution features. The final output is a 20×20 three-dimensional feature tensor with 256 channels, which serves as the input to the subsequent decomposition module. The decoder employs a 4-level upsampling structure, performing feature upsampling through bilinear interpolation. Simultaneously, it progressively fuses multi-scale feature maps from corresponding encoder levels to compensate for the loss of detailed features during downsampling, achieving spatial resolution restoration and ensuring the feature integrity of small-scale defects.

[0095] Specifically, the data processing diagram of the decomposition module is shown below. Figure 3 As shown, the decomposition module performs frequency domain low-rank decomposition and tensor singular spectrum decomposition on the spatial domain image features to obtain a fused feature tensor. The defect detection head is used to perform defect classification and bounding box regression on the fused feature tensor using anchor boxes to obtain defect detection and recognition results. A schematic diagram of its data processing process is shown below. Figure 4 As shown. The output layer is used to output the defect detection and identification results.

[0096] In this embodiment, the training method for the dam defect detection model includes:

[0097] The original multimodal images of the dam are acquired and preprocessed to obtain standardized image samples of the dam. Based on the standardized image samples of the dam, the standardized image samples are used as input features, the corresponding pre-labeled defect labels are used as output features, and the samples are divided into training set, validation set and test set according to a preset ratio.

[0098] A multidimensional hybrid loss function is constructed, which is a weighted sum of binary cross-entropy loss, Dice loss, perceptual loss, generalized cross-union ratio loss, tensor singular spectrum decomposition reconstruction loss and frequency domain low-rank decomposition reconstruction loss. Each loss term is configured with a weight hyperparameter to obtain the total loss function.

[0099] In this embodiment, the total loss function is expressed as:

[0100] ;

[0101] In the formula, This represents the loss value of the total loss function. Representing the binary cross-entropy loss respectively Dice loss Perceived loss Generalized intersection and comparison loss Tensor singular spectral decomposition reconstruction loss Frequency domain low-rank decomposition reconstruction loss The weighting coefficients, where, .

[0102] Based on the total loss function and the training set, the dam defect detection model is pre-trained, the weights of the backbone network are frozen, and the decomposition module and defect detection head are iteratively trained to obtain the pre-trained model.

[0103] Based on the pre-trained model, all network layers are unfrozen, and the dam defect detection model is jointly trained end-to-end using a pre-heated cosine annealing learning rate strategy to obtain a fine-tuned model.

[0104] Based on the fine-tuned trained model and the validation set, K-fold cross-validation is used to select the optimal hyperparameter combination to obtain the cross-validated model. Based on the cross-validated model, the test set is used for generalization validation, and incremental training is performed on standardized dam image samples with a recognition rate lower than a preset threshold to obtain the trained dam defect detection model.

[0105] Step 2.1: Perform layer-by-layer downsampling on the standardized image of the dam to extract shallow texture and deep semantics, and use the downsampling results as multi-scale features, with the deepest feature as the output feature; perform upsampling on the output feature, and fuse the multi-scale features to restore spatial resolution, outputting spatial domain image features in the form of a three-dimensional tensor.

[0106] Step 2.2: Perform frequency domain low-rank decomposition and tensor singular spectrum decomposition on the spatial domain image features to obtain the fused feature tensor.

[0107] Step 2.2.1: Perform Fourier basis expansion transformation on the spatial domain image features to the frequency dimension to obtain frequency domain features.

[0108] This embodiment obtains expanded spatial domain features by adding an extended window to the spatial domain image features; the boundaries of the expanded spatial domain features are constrained based on the window function to obtain windowed spatial domain features; Fourier basis expansion is performed on the windowed spatial domain features to map them to frequency amplitude spectrum and phase sensing basis functions to obtain frequency domain features.

[0109] This embodiment uses the Hanning extended window to add an extended window to the spatial domain image features. The formula for calculating the weight coefficients of the Hanning extended window is expressed as follows:

[0110] ;

[0111] In the formula, This represents the weighting coefficient of the Hanning extended window. Represents discrete sampling points. Represents the cosine function. This indicates the window length, with a value of 16.

[0112] In this embodiment, the frequency domain features are represented as:

[0113] ;

[0114] In the formula, Represents frequency domain characteristics, These represent the horizontal frequency index and the vertical frequency index in the frequency domain, respectively. Indicates the height of the spatial domain feature map. Indicates the width of the spatial domain feature map. Representing spatial domain image features at location Pixel feature value at that location This represents the weighting coefficient of the Hanning extended window. This represents the horizontal pixel coordinate index of the spatial domain feature map. This represents the vertical pixel coordinate index of the spatial domain feature map. Represents the natural constant. It represents the imaginary unit.

[0115] Step 2.2.2: Perform third-order Hankel tensor quantization and recombination on the frequency domain features to obtain a third-order tensor.

[0116] This embodiment obtains a multi-channel frequency domain feature map by acquiring the feature maps of each channel of the frequency domain feature map; it then acquires the neighborhood feature blocks of the multi-channel frequency domain feature map at each pixel location to obtain local frequency domain feature blocks; based on the local frequency domain feature blocks, it stacks them according to the feature channel dimension to construct a third-order Hankel matrix to obtain a third-order Hankel matrix; finally, it fuses the feature maps of each channel in the third-order Hankel matrix and reassembles them into a third-order Hankel frequency domain feature tensor.

[0117] Step 2.2.3: Perform tensor singular spectrum decomposition on the third-order tensor to obtain effective features.

[0118] This embodiment calculates the mean of each dimension of the global domain of the third-order tensor and subtracts the mean of the corresponding dimension from each element of the third-order tensor to obtain a centered third-order tensor. Modal expansion operations are then performed on the centered third-order tensor along the spatial dimension, neighborhood dimension, and feature channel dimension, decomposing it into three sets of two-dimensional matrices. Singular value decomposition is then performed on each of the three sets of two-dimensional matrices to obtain the left singular vector matrix, the singular value diagonal matrix, and the right singular vector matrix corresponding to each set of two-dimensional matrices. The singular values ​​in the singular value diagonal matrix are ordered according to their numerical values. Sort the data from largest to smallest; based on the noise distribution of the dam standardized image, a singular value threshold is preset, and singular values ​​in each singular value diagonal matrix that are greater than the singular value threshold, as well as the left and right singular vector matrices corresponding to the singular values, are retained; based on the retained singular values ​​in each singular value diagonal matrix that are greater than the singular value threshold, as well as the left and right singular vector matrices corresponding to the singular values, each set of two-dimensional feature matrices is reconstructed in reverse, and each set of two-dimensional feature matrices is reassembled according to the modal structure and dimensional order of the third-order tensor to obtain effective features.

[0119] In this embodiment, the centered third-order tensor is represented as:

[0120] ;

[0121] ;

[0122] In the formula, This represents the centered third-order tensor. Represents a third-order tensor. These represent the spatial neighborhood dimension index, the pixel sequence dimension index, and the feature channel dimension index of the third-order Hankel frequency domain feature tensor, respectively. Indicates the first The channel mean value calculated for all pixels in all spatial neighborhoods of the entire domain under each feature channel. It represents the total length of the spatial neighborhood dimension of a third-order tensor.

[0123] In this embodiment, the decomposition formula for performing singular value decomposition on the three sets of two-dimensional matrices is expressed as follows:

[0124] ;

[0125] In the formula, Represents the decomposition formula. Describes a left singular vector matrix. Represents a singular value diagonal matrix. Describes a right singular vector matrix. This indicates transpose.

[0126] Step 2.2.4: Use nonnegative matrix decomposition with cosine similarity regularization constraints and hyperplane projection constraints to divide the effective features and preset low-rank components of the features to obtain the decomposed features.

[0127] This embodiment obtains the effective features and preset low-rank components of the features as the input matrix to be decomposed; applies non-negativity constraints to the input matrix to obtain a non-negativity-constrained input matrix; based on the non-negativity-constrained input matrix, a weighted combination of the F-norm reconstruction loss term and the cosine similarity regularization term is used to construct a target loss function; based on the target loss function, non-negativity matrix decomposition is performed on the input matrix to obtain a coefficient matrix and a basis matrix; a hyperplane projection constraint is applied to the coefficient matrix and the basis matrix to obtain a coefficient matrix and a basis matrix that satisfy the hyperplane constraint; based on the coefficient matrix and the basis matrix that satisfy the hyperplane constraint, the decomposed features are reconstructed.

[0128] In this embodiment, the decomposition features are represented as:

[0129] ;

[0130] In the formula, Describes the decomposition features, where, For tensor reshaping operators, As an effective feature, The total dimension of the feature channels of the third-order tensor. For the preset low-rank components, The coefficient matrix that satisfies the hyperplane constraint is, where, , .

[0131] In this embodiment, the target loss function is expressed as:

[0132] ;

[0133] In the formula, This represents the loss value of the objective loss function. The loss term is reconstructed using the F-norm. For cosine similarity regularization, Denotes the square of the Frobenius norm of a matrix. This represents the function for calculating cosine similarity. This represents the weight coefficient of the cosine similarity regularization term.

[0134] In this embodiment, the cosine similarity regularization term is represented as:

[0135] ;

[0136] In the formula, Represents the preset low-rank components. The total number of basis vectors, Represents the preset low-rank components. The Column basis vectors Represents the preset low-rank components. The Column basis vectors This represents the L2 norm of a vector.

[0137] Step 2.2.5: Sparsify the decomposed features and classify the effects into layers according to the defect morphology, and output the fused feature tensor.

[0138] This embodiment obtains sparsified decomposition features by sparsifying the decomposition features; obtains the amplitude distribution of the sparsified decomposition features in the spatial frequency domain, and divides the amplitude distribution into low-frequency amplitude and high-frequency amplitude parts according to a preset spatial frequency threshold, obtaining low-frequency components and high-frequency components; based on the low-frequency components and the high-frequency components, the sparsified decomposition features are correspondingly divided into low-frequency defect features and high-frequency defect features, obtaining hierarchically classified features; the hierarchically classified features are fused to output a fused feature tensor.

[0139] Step 2.3: Use anchor frames to perform defect classification and bounding box regression on the fused feature tensor to obtain defect detection and recognition results.

[0140] Step 2.3.1: Based on the fused feature tensor, the defect location is predicted using anchor frames with a preset scale and aspect ratio to obtain the initial predicted bounding box.

[0141] In this embodiment, the aspect ratio of the anchor frame is set to 4:1, 2:1, 5:1, or 1:2 to adapt to the linear crack morphology.

[0142] Step 2.3.2: Based on the initial predicted bounding box, calculate the defect category confidence and bounding box coordinate offset of each anchor box to obtain the classification result and regression result.

[0143] Step 2.3.3: Based on the classification and regression results, use the generalized intersection-union loss function as the loss function for bounding box regression, calculate the loss value between the predicted bounding box and the true bounding box, and obtain the optimized bounding box regression result.

[0144] In this embodiment, the generalized intersection-union ratio loss function is expressed as:

[0145] ;

[0146] ;

[0147] ;

[0148] In the formula, Let GIoU represent the generalized intersection-union (CIU) loss function, and IoU represent the CIU loss. B represents the smallest bounding rectangle containing both the predicted and ground truth bounding boxes, and B represents the defect bounding box predicted by the dam defect detection model. The bounding box represents the actual annotation of the defect, ∩ represents the intersection, and ∪ represents the union. This indicates the area of ​​the bounding box of the region.

[0149] Step 2.3.4: Based on the classification results and the optimized bounding box regression results, output the defect detection and identification results.

[0150] Example 3

[0151] Based on the same inventive concept as other embodiments, this embodiment describes a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of the methods of Embodiment 1 or 2 described above.

[0152] Example 4

[0153] Based on the same inventive concept as other embodiments, this embodiment introduces a computer program product, including computer instructions that, when executed by a processor, implement the steps of the methods described in Embodiment 1 or 2 above.

[0154] In summary, this invention acquires infrared images of dam leakage and underwater dam structure images and performs preprocessing. After downsampling the standardized dam image layer by layer to extract shallow texture and deep semantics, it performs upsampling and fusion of multi-scale features, outputting spatial domain image features in the form of a three-dimensional tensor. The spatial domain image features are then subjected to frequency domain low-rank decomposition and tensor singular spectrum decomposition to obtain a fused feature tensor. This fused feature tensor is used for defect classification and bounding box regression, achieving unified feature modeling and cross-domain feature fusion for both infrared and underwater modal images. This effectively suppresses the masking of defect features by infrared clutter and water scattering noise, improving the detection accuracy of dam defects in complex underwater environments and solving the problem of dam defect features being easily masked by noise and inaccurately located in complex underwater environments.

[0155] This invention acquires infrared images of dam leakage and underwater dam structure images, performs preprocessing, and then uses an encoder in a backbone network to downsample and extract shallow textures and deep semantics. A decoder then upsamples and fuses multi-scale features to output three-dimensional tensor spatial domain image features. A decomposition module performs low-rank frequency domain decomposition and tensor singular spectrum decomposition on these spatial domain image features to obtain a fused feature tensor. This achieves unified feature modeling and cross-domain feature fusion for both infrared and underwater modal images in the spatial, frequency, and tensor domains. It effectively suppresses the masking of defect features by infrared clutter and water scattering noise, improving the detection accuracy of dam defects in complex underwater environments.

[0156] This invention transforms spatial domain image features into the frequency dimension by adding an extended window and performing Fourier basis expansion. The frequency domain features are then reconstructed using third-order Hankel tensor quantization to obtain a third-order tensor. The third-order tensor is then subjected to centering, modal expansion, singular value decomposition, filtering of effective singular values ​​based on a preset threshold for noise distribution in a dam-standardized image, and inverse reconstruction to obtain effective features after noise removal. This invention can separate defect features from the original image that is masked by noise, solving the problem of difficulty in extracting defect features in complex underwater environments.

[0157] This invention obtains decomposed features by performing nonnegative matrix decomposition with cosine similarity regularization constraints and hyperplane projection constraints on effective features and preset low-rank components of features. The decomposed features are then sparsified and divided into low-frequency defect features and high-frequency defect features according to spatial frequency thresholds before being fused to obtain a fused feature tensor. This enables hierarchical classification of trend defects and non-stationary fluctuation defects. The defect detection head uses anchor frames with preset scale and aspect ratio to classify defects on the fused feature tensor, and uses the generalized intersection-union loss function as the loss function for bounding box regression, thereby improving the bounding box positioning accuracy of slender cracks and micro-leakage.

[0158] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0159] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0160] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0161] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0162] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. A method for detecting and identifying defects in dams, characterized in that, include: The original multimodal images of the dam to be detected are acquired and preprocessed to obtain a standardized image of the dam. The original multimodal images of the dam include infrared images of dam leakage and underwater images of the dam structure. Based on the standardized image of the dam, the image is processed using a pre-trained dam defect detection model to output defect detection and identification results, including the type, location, and confidence level of the dam defect. The data processing method for the dam defect detection model includes: The standardized image of the dam is downsampled layer by layer to extract shallow texture and deep semantics, and the downsampling results are used as multi-scale features, with the deepest feature being used as the output feature. The output feature is then upsampled and multi-scale features are fused to restore spatial resolution, outputting spatial domain image features in the form of a three-dimensional tensor. The spatial domain image features are subjected to frequency domain low-rank decomposition and tensor singular spectrum decomposition to obtain the fused feature tensor: Anchor frames are used to classify defects and regress bounding boxes on the fused feature tensor to obtain defect detection and recognition results.

2. The dam defect detection and identification method according to claim 1, characterized in that, The network structure of the dam defect detection model includes: The input layer is used to receive the standardized image of the dam; The backbone network includes a multi-level stacked encoder and decoder. The encoder is used to downsample the dam-normalized image layer by layer to extract shallow texture and deep semantics and use the downsampling results as multi-scale features, with the deepest feature as the output feature. The decoder is used to upsample the output features and fuse the multi-scale features to restore spatial resolution and output spatial domain image features in the form of a three-dimensional tensor. The decomposition module is used to perform frequency domain low-rank decomposition and tensor singular spectrum decomposition on the spatial domain image features to obtain a fused feature tensor. The defect detection head is used to classify defects and regress bounding boxes on the fused feature tensor using anchor frames to obtain defect detection and recognition results. The output layer is used to output the defect detection and identification results.

3. The dam defect detection and identification method according to claim 1, characterized in that, The spatial domain image features are subjected to frequency domain low-rank decomposition and tensor singular spectrum decomposition to obtain a fused feature tensor, including: The spatial domain image features are subjected to Fourier basis expansion and transformed to the frequency dimension to obtain frequency domain features; The frequency domain features are recombined using third-order Hankel tensor quantization to obtain a third-order tensor; Perform tensor singular spectrum decomposition on the third-order tensor to obtain effective features; A nonnegative matrix decomposition with cosine similarity regularization constraints and hyperplane projection constraints is used to divide the effective features and the preset low-rank components of the features to obtain the decomposed features. The decomposed features are sparsified, and the effects are classified into layers according to the defect morphology, and the fused feature tensor is output.

4. The dam defect detection and identification method according to claim 3, characterized in that, The spatial domain image features are subjected to Fourier basis expansion and transformed to the frequency dimension to obtain frequency domain features, including: By adding an extended window to the spatial domain image features, the extended spatial domain features are obtained; Based on the window function constraint of the boundary of the extended spatial domain features, the windowed spatial domain features are obtained. The windowed spatial domain features are subjected to Fourier basis expansion operations to map them to frequency amplitude spectrum and phase sensing basis functions, thus obtaining frequency domain features.

5. The dam defect detection and identification method according to claim 3, characterized in that, The frequency domain features are reconstructed using third-order Hankel tensor quantization to obtain a third-order tensor, including: Obtain the feature maps of each channel of the frequency domain features to obtain a multi-channel frequency domain feature map; Obtain the neighborhood feature blocks of the multi-channel frequency domain feature map at each pixel location to obtain the frequency domain local feature blocks; Based on the frequency domain local feature blocks, a third-order Hankel matrix is ​​constructed by stacking them according to the feature channel dimension to obtain the third-order Hankel matrix. The feature maps of each channel in the third-order Hankel matrix are fused and recombined into a third-order Hankel frequency domain feature tensor.

6. The dam defect detection and identification method according to claim 3, characterized in that, Performing tensor singular spectrum decomposition on the third-order tensor yields effective features, including: The mean of each dimension is calculated over the entire domain of the third-order tensor, and the mean of the corresponding dimension is subtracted from each element in the third-order tensor to obtain the centered third-order tensor. The centered third-order tensor is subjected to modal expansion operations along the spatial dimension, neighborhood dimension, and feature channel dimension, respectively, and is decomposed into three sets of two-dimensional matrices; Singular value decomposition is performed on the three sets of two-dimensional matrices respectively to obtain the left singular vector matrix, singular value diagonal matrix and right singular vector matrix corresponding to each set of two-dimensional matrices, wherein the singular values ​​in the singular value diagonal matrix are sorted in descending order of value; Based on the noise distribution of the dam standardized image, a singular value threshold is preset, and singular values ​​in each singular value diagonal matrix that are greater than the singular value threshold, as well as the left singular vector matrix and the right singular vector matrix corresponding to the singular value are retained. Based on the singular values ​​greater than the singular value threshold in each retained singular value diagonal matrix, and the left and right singular vector matrices corresponding to the singular values, each set of two-dimensional feature matrices is reconstructed in reverse, and each set of two-dimensional feature matrices is reassembled according to the modal structure and dimensional order of the third-order tensor to obtain effective features.

7. The dam defect detection and identification method according to claim 3, characterized in that, A nonnegative matrix factorization with cosine similarity regularization and hyperplane projection constraints is used to divide the effective features and preset low-rank components to obtain the decomposed features, including: Obtain the effective features and the preset low-rank components of the features as the input matrix to be decomposed; Apply a nonnegativity constraint to the input matrix to obtain a nonnegativity-constrained input matrix. Based on the input matrix after the non-negativity constraint, the target loss function is constructed by weighted combination of F-norm reconstruction loss term and cosine similarity regularization term. Based on the target loss function, nonnegative matrix decomposition is performed on the input matrix to obtain the coefficient matrix and the basis matrix; Apply hyperplane projection constraints to the coefficient matrix and the basis matrix to obtain a coefficient matrix and a basis matrix that satisfy the hyperplane constraints; Based on the coefficient matrix and basis matrix that satisfy the hyperplane constraint, the decomposition features are reconstructed.

8. The dam defect detection and identification method according to claim 3, characterized in that, The decomposed features are sparsified, and the effects are stratified and classified according to the defect morphology. The fused feature tensor is then output, including: The decomposition features are sparsified to obtain the sparsified decomposition features. The amplitude distribution of the sparsified decomposition features in the spatial frequency domain is obtained, and the amplitude distribution is divided into a low-frequency amplitude part and a high-frequency amplitude part according to a preset spatial frequency threshold to obtain low-frequency components and high-frequency components. Based on the low-frequency components and the high-frequency components, the sparsified decomposition features are divided into low-frequency defect features and high-frequency defect features to obtain the hierarchical classification features. The features after hierarchical classification are fused to output a fused feature tensor.

9. The method for dam defect detection and identification according to claim 1, characterized in that, Anchor frames are used to classify defects and regress bounding boxes on the fused feature tensor, resulting in defect detection and recognition results, including: Based on the fused feature tensor, the defect location is predicted using anchor frames with preset scale and aspect ratio to obtain an initial predicted bounding box. Based on the initial predicted bounding box, the defect category confidence and bounding box coordinate offset of each anchor box are calculated to obtain the classification result and regression result; Based on the classification and regression results, the generalized intersection-union loss function is used as the loss function for bounding box regression. The loss value between the predicted bounding box and the true bounding box is calculated to obtain the optimized bounding box regression result. Based on the classification results and the optimized bounding box regression results, the defect detection and identification results are output.

10. The method for dam defect detection and identification according to claim 1, characterized in that, The training method for the dam defect detection model includes: The original multimodal images of the dam were acquired and preprocessed to obtain standardized image samples of the dam. Based on the standardized image samples of the dam, the standardized image samples are used as input features, the corresponding pre-labeled defect labels are used as output features, and the samples are divided into training set, validation set and test set according to a preset ratio. A multi-dimensional hybrid loss function is constructed, which is a weighted sum of binary cross-entropy loss, Dice loss, perceptual loss, generalized cross-union ratio loss, tensor singular spectrum decomposition reconstruction loss and frequency domain low-rank decomposition reconstruction loss. Each loss term is configured with a weight hyperparameter to obtain the total loss function. Based on the total loss function and the training set, the dam defect detection model is pre-trained, the weights of the backbone network are frozen, and the decomposition module and defect detection head are iteratively trained to obtain the pre-trained model. Based on the pre-trained model, all network layers are unfrozen, and the dam defect detection model is jointly trained end-to-end using a pre-warmed cosine annealing learning rate strategy to obtain a fine-tuned model. Based on the fine-tuned trained model and the validation set, K-fold cross-validation is used to select the optimal hyperparameter combination to obtain the cross-validated model. Based on the cross-validated model, generalization validation is performed using the test set, and incremental training is conducted on standardized dam image samples with recognition rates below a preset threshold to obtain a trained dam defect detection model.