A concrete dam crack area identification method based on detection-grid cooperation

CN122695486APending Publication Date: 2026-09-04HOHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611180495.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-05
Publication Date
2026-09-04

AI Technical Summary

Technical Problem

[0006]发明目的:针对常规YOLO网络检测裂缝时,检测框覆盖裂缝不完整、细枝末端漏检等问题,提供一种基于检测-网格协同的混凝土坝裂缝区域识别方法,通过引入稠密网格分类分支与检测框形成互补,同时无需额外人工标注,利用联合损失函数与纹理辅助的协同推理算法,实现裂缝区域的完整覆盖和精确定位,为结构安全评估提供支撑

Benefits of technology

本发明提出的方法有助于实现检测结果完整覆盖裂缝区域,通过40×40稠密网格分类预测,以连续网格填补检测矩形框无法覆盖的裂缝细枝和末端,经MaM算法纹理筛选后,裂缝区域覆盖率大幅提高。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122695486A_ABST
    Figure CN122695486A_ABST
Patent Text Reader

Abstract

The application discloses a concrete dam crack area identification method based on detection-grid cooperation, relates to the technical field of dam safety detection and deep learning, and aims to solve the technical defects of incomplete coverage, high missing detection rate and high training cost when a traditional YOLO series model adopts a boundary box regression mode to detect irregular cracks; the application adds a grid classification branch in a YOLO11 detection network, generates a grid training label dynamically by using an existing detection box label, realizes double-task cooperative training by designing a joint loss function, and adopts a texture-assisted cooperative reasoning algorithm to fuse the detection box and the grid classification result, so that the cracks that cannot be covered by the detection box are filled, and finally the crack position and quantity information facing vectorization evaluation are output; the application significantly improves the crack area coverage completeness and detection precision, while maintaining efficient detection speed, is suitable for the concrete dam crack detection task, and provides technical support for digital intelligent inspection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of dam safety monitoring and deep learning technology, specifically to a method for collaborative recognition of concrete dam crack target detection and grid classification using an improved convolutional neural network. Background Technology

[0002] Cracks often appear in key structures such as concrete dams, spillways, and water conveyance tunnels in water conservancy projects during long-term operation. Timely and accurate detection of the location and morphology of cracks is an important prerequisite for ensuring the safe operation of the project. However, due to the complex background and environment and many interferences, traditional manual detection is prone to misdetection and omission.

[0003] Currently, target detection methods based on the YOLO series are widely used for crack identification due to their high speed and accuracy. By combining them with UAV high-altitude inspection, crack identification and location can be achieved efficiently.

[0004] However, surface cracks in long-operation hydraulic engineering structures are generally long, thin, and irregular in shape, with significant variations in scale. A single crack may correspond to multiple predicted bounding boxes, leading to common problems such as missed detection of crack ends and fine branches, and background contamination within the bounding box, making it impossible to support quantitative assessments of crack quantity and coverage area. Furthermore, when using equipment such as drones for inspection, real-time requirements must be considered. With the increasing prevalence of drone inspection technology in hydraulic engineering, the large volume of real-world images of concrete dams places higher demands on the completeness, accuracy, and real-time performance of crack identification.

[0005] Therefore, the purpose of this invention is to overcome the shortcomings of traditional YOLO series models in crack detection, such as incomplete coverage, high false negative rate, and poor robustness, and to provide a method for identifying crack areas in concrete dams based on detection-grid collaboration. By adding a grid classification branch to the YOLO detection network, automatically generating grid training labels using existing detection boxes, and designing a joint loss and collaborative post-processing algorithm, complete coverage and accurate localization of crack areas are achieved. Summary of the Invention

[0006] Purpose of the invention: To address the problems of incomplete crack coverage and missed detection of fine branches when using conventional YOLO networks to detect cracks, this invention provides a method for identifying crack areas in concrete dams based on detection-mesh collaboration. By introducing dense mesh classification branches to complement the detection boxes, and without the need for additional manual annotation, this method utilizes a joint loss function and texture-assisted collaborative inference algorithm to achieve complete coverage and accurate localization of crack areas, thus providing support for structural safety assessment.

[0007] To achieve the above objectives, the technical solution adopted by this invention is: a method for identifying crack regions in concrete dams based on detection-grid collaboration, comprising the following steps: Step 1: Construct a dual-task detection model based on the YOLO11 network structure, which includes an object detection branch and a grid classification branch. The dual-task head accepts the first three feature layers for object detection and the fourth feature layer for grid classification, and then outputs bounding box prediction and dense grid classification prediction respectively.

[0008] Step 2: Dynamically generate grid ground truth labels. During the data loading stage, grid ground truth labels are dynamically generated based on the existing detection box annotations. The image is uniformly divided into a fixed number of grid units, and the intersection area ratio between each grid unit and the detection box is calculated and assigned a category. No additional manual annotation is required.

[0009] Step 3: Design a joint loss function, and optimize the detection loss and grid classification loss. The grid classification loss adopts a weighted combination of Dice Loss and Focal Loss, and sets high weights for positive samples to alleviate the class imbalance problem, while improving the recall rate to ensure the continuity of the gap.

[0010] Step 4: MaM (Matching and Masking) Collaborative Inference Filtering. During the inference phase, the texture-assisted MaM collaborative inference algorithm is used to perform secondary filtering on the mesh prediction. Positive meshes that intersect with the detection box are directly retained, while positive meshes that do not intersect are determined based on the local texture intensity score to determine whether to retain them, generating a corrected binary mesh map.

[0011] Step 5: Connected component merging and quantization output. The grids intersecting with the detection box area are also activated as positive sample grids. They are added to the grids corrected in Step 4 to obtain the total binary grid map. Then, the circumscribed rectangle is generated as the crack continuous location box through connected component analysis.

[0012] Furthermore, the specific steps for constructing the dual-task detection model in step 1 are as follows: 1-1 First, the input concrete dam crack image size is standardized to a fixed size. Then, features are extracted multiple times using the Conv and C3k2 modules to obtain multi-scale feature maps with output sizes of 320×320 for layer 0, 160×160 for layer 1, 80×80 for layer 3, 40×40 for layer 5, and 20×20 for layer 7. The number of channels is gradually increased with downsampling to extract disease features at different scales. These features are then fed into the neck fusion module to output a feature map of a certain size.

[0013] 1-2 The backbone network uses a YOLO11 configuration file with an adjusted output head, which can output an additional feature map for grid classification. The backbone consists of 10 core layers. Based on regular convolutions and C3k2, an SPPF module is introduced to expand the receptive field through multi-scale pooling, and a C2PSA module enhances the global context representation through cross-stage channels and spatial attention. The final output is a deep feature map rich in high-level semantics, ensuring discriminative power for slender cracks in complex backgrounds. After feature pyramid-path aggregation and bidirectional fusion in the neck layer, four feature layers adaptable to dual-task requirements are output, capturing the fine texture, medium-scale contour, and global distribution features of the cracks, while also meeting the fine feature requirements of the grid classification branch. The core convolution operation formula is as follows: (1)

[0014] In the formula, y i, j Let k be the (i, j)th pixel value of the convolutional output feature map, and w be the kernel size. m,n Let x be the weight at the (m,n)th position of the convolution kernel. i+m,,j+n is the pixel value at the corresponding position in the input feature map, and b is the convolution bias term. This formula achieves linear transformation and feature extraction of the input features.

[0015] The calculation formula for the channel attention mechanism in the C2PSA module is as follows: (2)

[0016] In the formula, σ is the Sigmoid activation function, MLP is a multilayer perceptron, GlobalAvgPool is a global average pooling operation, x is the input feature map, and Attention(C) is the channel attention weight of the output, which is used to adaptively weight the features of different channels, strengthen the disease feature channel, and suppress irrelevant background channels.

[0017] 1-3 The neck layer employs a bidirectional fusion mechanism of feature pyramid-path aggregation. From top to bottom, layer 10 is concatenated with layer 6 after 1×1 convolution and nearest-neighbor upsampling, then refined by the C3K2 module to obtain layer 14. Layer 14 is then concatenated with layer 4 after repeating the 1×1 convolution and upsampling operation, and finally refined by the C3K2 module to obtain layer 18. From bottom to top, layer 18 is downsampled by 3×3 convolution, concatenated with layer 16, and reconstructed by the C3K2 module to obtain layer 21. Layer 21 is then concatenated with layer 10 after repeating the downsampling operation, and finally reconstructed by the C3K2 module to obtain layer 24. The final output consists of feature layers 18, 21, and 24 for detection, and an additional feature layer 21 for grid classification.

[0018] 1-4 Design a dual-task detection head for both detection and grid classification, inheriting from the original YOLO11 detection head. It receives the four feature layers output from the neck and divides them into two branches to achieve collaborative detection and classification tasks: For the detection branch, the original YOLO11 Detect module is used, which receives the first three feature layers and outputs bounding box regression results and class predictions.

[0019] For the grid classification branch, the fourth feature layer is sequentially subjected to 1×1 convolution to adjust channels, adaptive average pooling to adjust the output, 3×3 convolution downsampling, batch normalization, non-linear activation function, and Dropout2D regularization. Finally, a binary classification grid prediction map of size [B, 2, 40, 40] is output by adjusting the number of channels through 1×1 convolution, where B is the number of images processed in a single batch. The grid branch weights are initialized using the Kaiming normal distribution. The weights of positive sample pixels in grid classification are increased, and the biases of the last convolution layer of the grid branch are differentially initialized, with the negative class background bias initially set to positive and the positive class crack bias initially set to negative. The 1×1 convolution formula for the grid classification branch is as follows, used to achieve channel dimension compression and mapping: (3)

[0020] In the formula, W is a 1×1 convolution kernel weight matrix, and the parameter dimension is [C]. in C out [, 1, 1], b is the convolution bias, and y is the convolution output feature map. This formula can flexibly adjust the number of channels without changing the feature map spatial size, thereby reducing the computational cost of the model.

[0021] During training, the network returns a pair of detection feature lists and grid predictions. During inference, it outputs a pair of decoded detection boxes, intermediate features, and grid probability maps for subsequent collaborative post-processing.

[0022] 1-5 Differentiated weight initialization is adopted for different layers of the network, focusing on optimizing the convergence effect of the grid classification branch; gradient vanishing is alleviated by initializing the convolutional layer with a normal distribution, non-linear activation function, BatchNorm2d batch normalization, and differential bias of the grid classification output layer, ensuring that training works properly and accelerating feature extraction and network convergence.

[0023] The core hyperparameters of the model (1-6) include the number of detection categories (nc) and the number of grid classification categories (nc). grid Grid classification enabled flag (use) grid Grid size size depth multiple width multiple multiple .

[0024] The training dataset (1-7) consists of a mix of images taken on-site at concrete dams and publicly available crack databases, covering multiple types of defects and including various lighting, noise, and concrete texture backgrounds to enhance robustness. All images only provide crack detection bounding boxes; ground truth labels for the grid cells are dynamically generated based on the intersection of the detection boxes and grid cells during data loading. The dataset is randomly divided into training, validation, and test sets according to a certain ratio.

[0025] The optimizer used in sections 1-8 is the Adam optimizer. By setting a small initial learning rate and enabling cosine annealing to decay the learning rate, the model training and convergence are stabilized, and overfitting is prevented by weight decay.

[0026] 1-9 Set larger single training batches to ensure normalization stability during training, enable early stopping tolerance to prevent overfitting due to overtraining, and simultaneously evaluate and record metrics on the validation set during training. Disable mixed precision training, enable non-maximum suppression, and set an appropriate confidence threshold to ensure the reliability of the final detection results.

[0027] 1-10 Employing multi-strategy data augmentation enhances the model's generalization ability to simulate the variable distribution of diseases in complex engineering environments, while effectively suppressing overfitting.

[0028] Furthermore, the specific steps for generating the dynamic mesh labels in step 2 are as follows: 2-1 By adding configuration items to the YOLODataset dataset class, the grid label generation function can be enabled and the grid generation density can be controlled to adapt to different detection needs.

[0029] 2-2 The input image is uniformly divided into 40×40 seamless, non-overlapping grid cells. Since the YOLO data annotation system uses normalized coordinates (i.e., the image width and height are normalized to the [0,1] interval), the coordinates of any point in the image are represented by the ratio of its actual pixel position to the image width and height. Under this normalized coordinate system, the size of each grid cell is w. cell ×h cell w cell =1 / 40, h cell =1 / 40, where 40 is the grid density of the image in both the vertical and horizontal directions.

[0030] 2-3 Read the normalized detection box center position x from the standard YOLO annotation file center , y center Width and height are represented as (x, y ... center , y center (, width, height) and corresponding categories, with the crack category index number being 0.

[0031] 2-4 Traverse each grid cell and calculate its IoA with all detection boxes, i.e., intersection area / grid area. Assign crack or background categories according to the preset label threshold. If the maximum IoA is greater than the label threshold, mark it as crack with a corresponding mapping value of 1. Otherwise, mark it as background with a corresponding mapping value of 0. Generate the full-image grid label matrix.

[0032] 2-5 In the data loader, the grid labels of each sample are stacked into a tensor of [B, 40, 40] and input into the model along with the detection data. The grid labels are generated during the data loading process and can be calculated in real time based on the transformed detection boxes after image data augmentation and validation preprocessing. Therefore, the grid labels automatically maintain complete consistency with the detection boxes in terms of displacement, scaling, flipping, and cropping changes, without the need for additional alignment operations, ensuring the accuracy and consistency of the grid supervision signal during the training and validation phases.

[0033] Furthermore, the specific steps for designing the joint loss function in step 3 are as follows: 3-1 Total Loss Formula: (4)

[0034] In the formula L det It is to detect loss, L grid For grid classification loss, λ grid The weights are used for grid classification loss.

[0035] 3-2 Detection Loss L det The original v8DetectionLoss loss function of YOLO11 is adopted, which includes bounding box regression loss, class loss and distribution focus loss.

[0036] 3-3 Grid Classification Loss L grid A weighted combination of Dice Loss and Focal Loss is used, as shown in the formula: (5)

[0037] The Dice Loss function is expressed as follows: (6)

[0038] In the formula, ε = 1 × 10 -6 w represents the pixel weight, with a foreground pixel weight of 7.0 and a background pixel weight of 0.5. p represents the grid prediction probability value, and t represents the grid's true label.

[0039] The Focal Loss function is expressed as follows: (7)

[0040] In the formula, α t =0.6, γ=2.0, P t The predicted probability of the correct category.

[0041] 3-4 After the network output is parsed and separated, the detection feature list is fed into v8DetectionLoss to calculate the three components of the detection loss and the total detection loss L. det The grid prediction tensor, after being processed by softmax to obtain the positive class probability, is fed into GridDiceFocalLoss along with the grid labels to calculate L. grid The final total loss L is obtained. total During training, each loss component is recorded separately, resulting in a 4-dimensional loss tensor.

[0042] Furthermore, the specific steps of the MaM collaborative reasoning algorithm in step 4 are as follows: 4-1 Read the model output, process the grid predictions using softmax, and then extract the positive class probability to obtain the grid score map. scores Grids that are predicted to be positive and whose probability is greater than the confidence threshold are marked as positive grids.

[0043] 4-2 For each positive grid, broadcast the calculation of its intersection IoA with all detection boxes. If IoA > 0 and the detection box category is crack, then directly retain it as a crack grid.

[0044] 4-3 For positive grids not covered by the detection box, extract the corresponding image patch in the original image and calculate the local texture intensity score: (8)

[0045] In the formula, GM is the average gradient magnitude calculated by the Sobel operator. If the score is higher than the texture threshold, it is retained as a cracked mesh; otherwise, it is suppressed as background.

[0046] 4-4 After two rounds of filtering, the retained positive grids form the corrected binary grid diagram.

[0047] Furthermore, the specific steps for connected component merging and quantization output in step 5 are as follows: 5-1 Merge the binary mesh image corrected by MaM with all the detection box occupancy areas. After the detection boxes are normalized and mapped to a 40×40 grid space, mark all grid cells that intersect with the detection boxes as 1 to generate the merged binary mesh image.

[0048] 5-2 The 8-neighborhood connected component labeling algorithm is used to perform connected component analysis on the merged binary mesh graph, and an identifier is assigned to each independent region starting from 1.

[0049] 5-3 Post-process the connected component labeling results, retaining connected regions that contain at least one detection box; merge and reorder overlapping connected regions to obtain the final crack connected regions.

[0050] 5-4 Generate a bounding rectangle for each independent connected region, calculate its row and column boundaries in the 40×40 grid space, and then convert them to the pixel coordinates of the model input image as the final crack location box.

[0051] The present invention provides a method for identifying crack regions in concrete dams based on detection-grid collaboration, which includes a dual-task detection model with a target detection branch and a grid classification branch. By combining the bounding box and grid detection results, MaM collaborative inference filtering is performed to obtain a fine grid division and a complete detection bounding box for the crack region.

[0052] The proposed dual-task detection model adds a grid classification head to the original YOLO11 detection head. By separating the data structures, the two task branches can run in parallel efficiently without interfering with each other, giving full play to the speed and accuracy advantages of YOLO11.

[0053] The dynamic grid label generation mechanism described above can generate labels quickly and accurately based on the bounding box annotations, avoiding the heavy burden of dense grid annotations, and can adaptively enhance transformations to improve the robustness of grid classification.

[0054] The proposed joint loss function retains the original YOLO11 loss function while adding a loss function suitable for grid classification. This ensures stable training and optimization for detection, and allows for balanced and stable convergence of the two tasks by adjusting the grid classification weights.

[0055] The MaM collaborative reasoning screening, after comprehensive detection and grid output, not only ensures the accuracy advantage of detection, but also fully leverages the high recall advantage of complete grid coverage. Finally, it removes misjudged grids through simple and rapid grid texture filtering, thereby improving accuracy.

[0056] Compared with the prior art, the present invention has the following beneficial effects: The method proposed in this invention helps to achieve complete coverage of the crack area by using a 40×40 dense grid for classification and prediction. The continuous grid fills in the crack branches and ends that cannot be covered by the detection rectangle. After texture filtering by the MaM algorithm, the crack area coverage is greatly improved.

[0057] In this invention, the grid labels are dynamically generated from the existing detection box annotations, eliminating the need for additional dense classification annotations.

[0058] The joint loss function designed in this invention can effectively suppress the inundation of the loss by the background mesh, helping to maintain a low false negative rate in complex concrete dam backgrounds. The mesh branch only adds about 0.3M parameters, and the overall inference speed and detection performance are basically on par with the original YOLO11 model. Moreover, the output structure is more in line with the crack morphology of concrete dams, providing technical support for efficient UAV inspections. Attached Figure Description

[0059] Figure 1 This is a schematic diagram of the concrete dam crack area identification method based on detection-grid collaboration in this embodiment.

[0060] Figure 2 This is a schematic diagram of the dual-task detection model architecture in this embodiment; Figure 3 This is a schematic diagram of the network structure in this embodiment; Figure 4 This is a schematic diagram of the dynamic grid label generation process in this embodiment; Figure 5 This is a schematic diagram of the joint loss calculation process in this embodiment; Figure 6 This is a schematic diagram of the MaM collaborative reasoning process in this embodiment; Figure 7 This is a schematic diagram comparing the crack area recognition results in this embodiment; Figure 8 This is a schematic diagram illustrating the dynamic label generation effect in this embodiment. Detailed Implementation

[0061] The present invention will be further illustrated below with specific implementation examples. These implementation examples are only for illustrating the present invention and are not intended to limit the scope of the present invention.

[0062] This embodiment provides a method for identifying crack areas in concrete dams based on detection-grid collaboration. The specific process is as follows: Figure 1 As shown, the process includes data preprocessing, model building and training, and prediction postprocessing. The specific implementation flow includes: Step 1, constructing a dual-task detection model, such as... Figure 2 As shown.

[0063] Specifically, step 1 includes: 1-1 Network structure diagram as follows Figure 3As shown, the input concrete dam crack image is standardized to 640×640 pixels. The first layer uses a 3×3 convolution with a stride of 2, outputting 64 channels, encoding the image into a basic feature map of 320×320×64 pixels. Features are extracted alternately using the Conv and C3k2 modules multiple times, resulting in multi-scale feature maps with output sizes of 320×320 (layer 0), 160×160 (layer 1), 80×80 (layer 3), 40×40 (layer 5), and 20×20 (layer 7). The number of channels gradually increases with downsampling to extract features of the damage at different scales. The shallow 80×80 feature layer has a low downsampling factor and rich details, which is specifically used to capture minute cracks and micro-cracks, accurately locate crack edges and textures, and avoid the loss of subtle features; the middle 40×40 feature layer has a moderate downsampling factor, which takes into account both details and global semantics, captures medium-scale cracks, and provides a feature interaction bridge for dual task branches; the deep 20×20 feature layer has a high downsampling factor, which captures cracks of globally distributed size.

[0064] 1-2 Based on the original YOLO11, the output head configuration is adjusted to output an additional feature map for grid classification, with a total of 10 core layers in the backbone. In addition to regular convolution and C3k2, an SPPF module is introduced to expand the receptive field through multi-scale pooling, and a C2PSA module enhances the global context representation through cross-stage channel and spatial attention. The final output is a deep feature map rich in high-level semantics, ensuring discriminative power for slender cracks in complex backgrounds. After completing the bidirectional fusion of feature pyramid and path aggregation in the neck section, four feature layers adapted to dual-task requirements are output, with the scale, number of channels, and functional localization of each feature layer strictly matching the subsequent branches. The first three feature layers have sizes of 80×80, 40×40, and 20×20, corresponding to downsampling factors of 8, 16, and 32, and channel numbers of 256, 512, and 1024, respectively, capturing the fine texture, medium-scale contour, and global distribution features of the cracks. The fourth feature layer has a size of 40×40, which retains more crack details, adapts to the fine feature requirements of the grid classification branch, and is also compatible with the grid density, achieving denser recognition of crack regions. The number of channels is determined by depth. multiple =0.5, width multiple =0.5 Lightweight setting.

[0065] 1-3 The neck layer employs a bidirectional fusion path of feature pyramid-path aggregation. From top to bottom, the 10th layer with 1024 channels is adjusted by 1×1 convolution, upsampled by 2 times the nearest neighbor, and then concatenated with the 6th layer. It is then refined by the C3K2 module to obtain the enhanced 512-channel 14th layer. The 14th layer is then subjected to the same 1×1 convolution and upsampling operation and concatenated with the 4th layer. It is then passed through the C3K2 module to obtain the 18th layer. From bottom to top, the 18th layer is downsampled by 3×3 convolution with a stride of 2 and the number of channels is adjusted. It is then concatenated with the 16th layer and reconstructed by the C3K2 module to obtain the 21st layer. The 21st layer is then subjected to the same downsampling operation and concatenated with the 10th layer and reconstructed by the C3K2 module to obtain the 24th layer. Finally, three feature layers for detection (18th, 21st, and 24th layers) and an additional 21st feature layer for grid classification are output. Bidirectional fusion allows for full interaction between shallow details and deep semantics, taking into account both local information and a larger receptive field, thus improving the detection accuracy of slender cracks. A feature pyramid-path aggregation bidirectional fusion mechanism is employed for the neck region.

[0066] The dual-task head (1-4) is implemented using a custom `DetectAndPatchClassify` class, which inherits from the original YOLO11 `Detect` class and combines detection and localization with grid classification functionality. This class requires the number of detection categories (`nc`), the feature map channel tuple (`ch`), and the grid size for initialization. The first three channels of `ch` are used for initializing the parent class `Detect`, while the fourth channel is used to construct a separate grid classification branch, enabling independent construction and collaborative optimization of the two tasks.

[0067] (1) Detection branch settings

[0068] The detection branch adopts the core structure of the original YOLO11 Detect module, optimized to meet the single-category detection requirements of concrete dam cracks. It is specifically responsible for bounding box regression and category determination of crack targets. This branch receives the first three feature layers from the neck output, with sizes of 80×80, 40×40, and 20×20 respectively. It utilizes multi-scale feature fusion and complementarity to achieve full-scale coverage detection from fine cracks to large-area cracks. The number of detection categories nc=1; the maximum bounding box regression value reg. max =16, improves the positioning accuracy of bounding boxes for slender and irregular cracks, and reduces offset problems.

[0069] (2) Grid classification branch settings

[0070] The grid classification branch specifically refines the fourth feature layer output from the neck region. Through a lightweight process, the features are mapped to a binary classification grid prediction map, achieving dense coverage recognition of the crack region. The specific operation sequence is as follows: By using a 1×1 convolution operation to adaptively match the number of input channels to the fourth feature layer, and fixing the number of output channels to 64, without setting a bias term, channels are compressed and features are fused without changing the feature map space size, reducing the amount of computation and preserving the key details of the crack. We employ BatchNorm2d batch normalization, which is consistent with the normalization of the detection branch, to standardize feature extraction, accelerate training convergence, prevent gradient vanishing or exploding problems, and improve model stability. The SiLU activation function enhances the nonlinear representation of features, alleviates gradient vanishing, and better distinguishes cracks from the background. Adaptive average pooling is used to fix the output feature map size at 40×40, ensuring that even when the size of the intermediate feature map changes slightly, the output feature map can still match the preset grid, providing support for a unified spatial dimension. The output is 32 channels through a 3×3 convolution operation. Padding=1 is set to ensure that the spatial size remains unchanged. No bias term is set to further refine the crack features and suppress background noise. Dropout2D regularization is used to randomly discard some feature channels, reducing overfitting, improving the model's generalization ability in complex backgrounds, and adapting to different lighting and noise environments. Finally, the 32 channels are mapped to 2 channels through 1×1 convolution, and the binary classification output is a grid prediction map of size [B, 2, 40, 40], where B is the number of images processed in a single batch, and the 2 channels correspond to the prediction probabilities of the background and cracks, respectively.

[0071] During forward propagation, the length of the input feature map list x is fixed at 4, and the output logic is designed differently according to the training / inference modes: Training mode return (det) output , pc output ) binary tuple. det output The first three feature maps are output by the detection head as an undecoded list of raw features, which is used to calculate the detection branch loss; pc output It is a 40×40×2 grid prediction tensor used to calculate the grid classification loss. The binary output ensures that the two losses are calculated independently and optimized together.

[0072] Inference mode returns triples (pred) boxes , pred feats , pc output ). pred boxes These are the detection boxes after decoding and confidence filtering, which can be directly used for post-prediction processing; feats To detect intermediate feature maps of branches, for visualization or to assist post-processing; pc outputThe predicted grid score, normalized by softmax, is used for subsequent MaM collaborative inference to improve the completeness and accuracy of crack identification.

[0073] 1-5 Differentiated weight initialization is used for different layers of the network, with a focus on optimizing the grid classification branch: (1) All convolutional layers are initialized with the Kaiming normal distribution and adapted to the SiLU activation function to alleviate gradient vanishing and facilitate the initial capture of crack features; (2) All BatchNorm2d layer weights are initialized to 1 and biases to 0 to ensure that regularization plays a normal role in the early stage of training, and to accelerate feature standardization extraction and network convergence; (3) The last layer of the grid classification branch uses a 2-channel 1×1 convolution with differential bias settings. The negative class background bias is 2.0, and the positive class crack bias is -1.0. By combining the grid classification loss with a positive sample weight of 7.0, the grid branch output is biased towards the background in the early stage of training, generating a larger positive sample loss backpropagation gradient, driving the grid branch to quickly focus on crack features, and avoiding convergence lag or unidirectional background optimization.

[0074] 1-6 The core hyperparameter configuration of the model includes the number of detection categories nc=1, and the number of grid classification categories nc. grid =2, use grid =True, gird size =40; depth multiple multiple =0.5, width multiple multiple =0.5.

[0075] The training dataset (1-7) consists of a mix of images taken on-site at concrete dams and publicly available crack databases, covering five types of defects: cracks, spalling, exudates, leakage, and spalling. Images are uniformly in PNG format and include various lighting, noise, and concrete texture backgrounds to enhance robustness. All images only provide crack detection bounding boxes; ground truth labels for the grid cells are dynamically generated during data loading from the intersection of the detection boxes and grid cells. The dataset is randomly divided into training, validation, and test sets in a 7:2:1 ratio.

[0076] The optimizer used in steps 1-8 is the Adam optimizer, with an initial learning rate of lr0=0.001. Cosine annealing is used to decay the learning rate; momentum is 0.937, weight decay is 0.01, and the learning rate is warmed up for 10 rounds with a warm-up momentum of 0.8 and a warm-up bias learning rate of 0.1. In later training rounds, the learning rate is slowly reduced to 10% of the initial learning rate to ensure stable convergence of the model.

[0077] 1-9 A single training batch size of 15 is used, with a total of 150 training epochs; early stopping tolerance patience is 70, and evaluation is performed simultaneously on the validation set during training. Mixed precision training is disabled, non-maximum suppression is enabled, and the confidence threshold is set to 0.2 to ensure the reliability of the final detection results.

[0078] 1-10 Multi-strategy data augmentation is adopted to improve the model's generalization ability. The enhancement coefficients for hue, saturation, and brightness are 0.015, 0.7, and 0.4, respectively. Random rotation ±5°, translation 0.1, scaling 0.1, clipping 2°, and perspective transformation 0.001 are used. The horizontal and vertical flip probabilities are both 0.2. Mosaic, mixup, and copy-paste enhancements are enabled with a probability of 0.2 to simulate the variable distribution of defects in complex engineering environments and effectively suppress overfitting.

[0079] Step 2, the dynamic mesh label generation process, as follows: Figure 3 As shown.

[0080] Specifically, step 2 includes: 2-1 In the YOLODataset dataset class, use the dataset configuration item "use" grid and grid size Controls the enabling of the grid label generation function and the grid generation density. In this embodiment, the grid... size =40.

[0081] 2-2 The input image is uniformly divided into 40×40 grid cells with no gaps or overlaps. After preprocessing with LetterBox and other methods, the image is normalized to the model input size, with each grid cell having a normalized size of w. cell ×h cell , where w cell =1 / 40, h cell =1 / 40.

[0082] 2-3 Read the normalized detection box coordinates (x, y) from the dataset annotations center , y center (width, height) and category information in tensor form for the bbox classes The crack category index is 0 in the YOLO annotation.

[0083] 2-4 For each grid cell (g) x , g y This process iterates through all detection boxes in the image and calculates the ratio IoA (intersection area between the grid cell and the detection box) to the area of ​​the grid cell. Specifically, the grid cell boundary is [g...]. x ×w cell , g y ×hcell ,(g x+1 )×w cell , (g y+1 )×h cell The detection box boundary is defined by (x) center , y center The coordinates (x1, y1, x2, y2) of the rectangle are converted to the coordinates of the top-left and bottom-right corners of the rectangle, respectively. This facilitates the calculation of the intersection area. If the maximum IoA is greater than the preset label threshold of 0.25, the grid is marked as a crack category and mapped to 1; otherwise, it is marked as background and mapped to 0. When there are no detection boxes in the image, all grids are marked as background.

[0084] 2-5 In the data loader, the grid labels of each sample are stacked into a tensor of [B, 40, 40], where B is the number of images processed in a single batch. This tensor is then used as the value of the gridlabels key in the batch data and input into the model along with the detection data such as img, cls, and bboxes for training. In subsequent enhancements, the grid labels are updated synchronously with the detection box labels after the image changes, without the need for additional alignment operations.

[0085] Step 3, calculate the joint loss function as follows: Figure 4 As shown.

[0086] Specifically, step 3 includes: 3-1 The total loss is obtained by weighted fusion of the detection branch loss and the grid classification branch loss, as shown in the following formula: (9)

[0087] In the formula L det It is to detect loss, L grid For grid classification loss, λ grid The weight for the grid classification loss is set to 1.2 in this paper to balance the loss ratio between the two tasks and ensure that the grid classification task and the detection task are optimized in a coordinated manner.

[0088] 3-2 Detection Loss L det The v8DetectionLoss method is used, which consists of a weighted average of three parts: bounding box regression loss, class loss, and distribution focus loss, as detailed below: (10)

[0089] In the formula L CIoU L is the bounding box regression loss, with a weight of 3.5, used to accurately measure the overlap and positional deviation between the predicted and ground truth bounding boxes; BCE This is a category loss function with a weight of 0.2, used to achieve accurate classification of disease categories; L DFLThe distributed focus loss, with a weight of 0.5, is used to optimize the prediction accuracy of bounding box coordinates and alleviate the class imbalance problem.

[0090] 3-3 The grid classification loss adopts a weighted combination of Dice Loss and Focal Loss, which takes into account both the integrity of the diseased area and the optimization of hard-to-classify pixels. The formula is as follows: (11)

[0091] In the formula, the weights of the two losses are uniformly set to 1.5 to ensure that Dice Loss and Focal Loss work together to guarantee the classification accuracy of grid images and enhance the learning of hard-to-classify pixels.

[0092] (1) The Dice Loss formula is used to measure the degree of overlap between the grid prediction results and the true labels, and to alleviate the problem of imbalance between positive and negative samples. The formula is as follows: (12)

[0093] In the formula, ε = 1 × 10 -6 To avoid the case where the denominator is 0, w is the pixel weight, where the foreground pixel weight is set to 7.0 and the background pixel weight is set to 0.5 to highlight the learning priority of the foreground disease area; p is the grid prediction probability value and t is the grid true label.

[0094] (2) The Focal Loss formula is used to focus on pixels that are difficult to classify and reduce the loss weight of background pixels that are easy to classify. The formula is as follows: (13)

[0095] In the formula, α t To balance the loss between positive and negative samples, γ is set to 0.6. γ is used to adjust the loss weight for hard-to-classify samples and is set to 2.0. P t To predict the correct category probability, the loss signal of positive samples (crack mesh) is further amplified, driving the model to focus on learning the features of the crack region.

[0096] 3-4 After the network output is parsed, the list of detected features is fed into v8DetectionLoss to calculate L. det And the detection loss has three components, and the bounding box loss L CIoU Classification loss L BCE Distribution focal loss L DFL The grid prediction tensor [B, 2, 40, 40] is processed by softmax to obtain the positive class probability, and then fed into GridDiceFocalLoss along with the grid labels to calculate L. grid Then L total = Ldet +1.2 × L grid Calculate the total loss. During training, record a 4-dimensional tensor [box] for each loss component. loss cls loss ,dfl loss grid loss ].

[0097] Step 4, MaM collaborative reasoning process, such as Figure 6 As shown.

[0098] Specifically, step 4 includes: 4-1 The model inference outputs a grid prediction tensor [B, 2, 40, 40] and a list of detection boxes. The grid predictions are processed using softmax, and the positive class probabilities are then calculated to obtain the grid score map. scores The grid has a shape of [B, 40, 40]. scores Grids with a confidence level greater than 0.5 of the grid classification threshold are marked as positive grids.

[0099] 4-2 For each positive grid (g) x , g y Calculate the area ratio (IoA) of the intersection between the normalized boundary and all detection boxes. If the detection box is in pixel coordinates, it needs to be normalized to the [0,1] interval first. If max(IoA) > 0, then the mesh intersects with the detection box and is directly retained as the crack mesh.

[0100] 4-3 For positive grids not covered by any detection boxes, their texture features need to be verified on the original image. The model input image size is 640×640, and the number of grids is 40×40, so each grid corresponds to an image region size of 16×16 pixels. Let the grid index be (g... x , g y ), corresponding pixel region: x1 = g x × 16, x2 = (g x + 1) × 16, y1 = g y × 16,y2 = (g y + 1) × 16. Convert the region in the model input tensor to a grayscale image, calculate the Sobel gradients in the x and y directions for each image patch, take the gradient magnitudes, calculate the mean, and normalize the score: (14)

[0101] In the formula, GM is the average gradient magnitude calculated by the Sobel operator. If the score > 0.3, the corresponding region of the mesh is determined to be a crack branch or end, and is retained as a crack mesh; otherwise, it is determined to be noise or non-crack texture (such as stains or shadows), and is suppressed as background.

[0102] 4-4 Corrected Mesh Generation: After two rounds of filtering, all the retained positive meshes form the corrected binary mesh Mam_grid with shape [40, 40], crack = 1, and background = 0.

[0103] Step 5: Connected component merging and quantization output

[0104] Specifically, step 5 includes: 5-1 The binary mesh image corrected by MaM is merged with the grid area occupied by the detection box. After normalization, the detection box is mapped to a 40×40 grid space, and all grid cells intersecting with the detection box are marked as 1, generating the merged binary mesh image.

[0105] 5-2 The merged binary mesh graph M merged An 8-neighborhood connected region labeling algorithm is used to assign a unique identifier starting from 1 to each independent connected region, and the total number of connected regions is counted.

[0106] 5-3 Post-process the connected component labeling results, filter out connected regions that do not intersect with detection boxes after merging, and retain connected regions that contain at least one detection box; merge the overlapping connected regions and recount them to obtain the final crack connected regions.

[0107] 5-4 For each final crack-connected region, calculate its row and column boundaries in a 40×40 grid space [x min ,x max , y min , y max ], converted to the model input image pixel coordinates: x1 = x min × 16, y1 = y min × 16;x2 =(x max +1) × 16, y2 = (y max +1) × 16. The coordinates of this rectangle are the final location frame of the crack area.

[0108] The proposed dual-task detection head adds a grid classification branch to the original YOLO11 detection head, using a 40×40 dense grid for classification prediction to compensate for the incomplete coverage of the rectangular detection box. A comparison of the crack region recognition performance with the native YOLO11 in this example is shown in the figure below. Figure 7 As shown.

[0109] The proposed dynamic grid label generation mechanism can obtain grid training signals without additional manual annotation and can adaptively enhance transformations. The grid label generation effect is shown in the image below. Figure 8 As shown.

[0110] The proposed Dice-Focal joint loss can effectively address the severe imbalance between positive and negative samples in grid classification, while stabilizing dual-task training by adjusting the grid classification weights.

[0111] The proposed texture-assisted MaM collaborative inference algorithm comprehensively considers the detection and mesh classification outputs and introduces extremely lightweight texture filtering. While maintaining high recall, it suppresses false detections of non-crack textures, and can obtain high-precision mesh regions and complete detection boxes for cracks.

[0112] This method proposes a new basic detection framework, which is applicable to the complete detection of cracks in different concrete dams and has a certain degree of universality.

[0113] The above description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be modified within the scope of the concept described herein by means of the above teachings or the technology or knowledge in related fields.

Claims

1. A method for identifying crack regions in concrete dams based on detection-grid collaboration, characterized in that, Includes the following steps: Step 1: Construct a dual-task detection model based on the YOLO11 network structure, which includes an object detection branch and a grid classification branch. The dual-task head receives the first three feature layers for object detection and the fourth feature layer for grid classification, and then outputs bounding box prediction and dense grid classification prediction respectively. Step 2: Dynamically generate grid ground truth labels. During the data loading stage, grid ground truth labels are dynamically generated based on the existing detection box annotations. The image is uniformly divided into a fixed number of grid units, and the intersection area ratio between each grid unit and the detection box is calculated and assigned a category. Step 3: Design a joint loss function, and simultaneously optimize the detection loss and grid classification loss. The grid classification loss adopts a weighted combination of Dice Loss and Focal Loss, and sets high weights for positive samples. Step 4: MaM Collaborative Inference Filtering. During the inference stage, the texture-assisted MaM collaborative inference algorithm is used to perform secondary filtering on the mesh prediction. Positive meshes that intersect with the detection box are directly retained, while positive meshes that do not intersect are determined based on the local texture intensity score to determine whether to retain them, generating a corrected binary mesh map. Step 5: Connected component merging and quantization output. The grids intersecting with the detection box area are also activated as positive sample grids. They are added to the grids corrected in Step 4 to obtain the total binary grid map. Then, the circumscribed rectangle is generated as the crack continuous location box through connected component analysis.

2. The method for identifying crack regions in concrete dams based on detection-grid collaboration according to claim 1, characterized in that, The construction of the dual-task detection model in step S1 specifically includes: After unifying the size of the input concrete dam crack images, multi-scale features are extracted through the backbone network. The backbone network introduces the SPPF module to expand the receptive field through multi-scale pooling on the basis of Conv and C3k2 modules, and introduces the C2PSA module to enhance the global context representation through cross-stage channels and spatial attention. The neck uses a feature pyramid-path aggregation bidirectional fusion mechanism to fuse the multi-scale features output by the backbone network, and outputs four feature layers adapted to dual task requirements. The first three feature layers are input to the target detection branch, and the fourth 40×40 feature layer is input to the grid classification branch. The grid classification branch sequentially processes the fourth feature layer of the input through 1×1 convolution to adjust the channels, batch normalization, SiLU activation, adaptive average pooling, 3×3 convolution downsampling, and Dropout2D regularization. Then, it outputs a binary classification grid prediction map with a size of [B,2,40,40] through 1×1 convolution, where B is the number of images in a single batch, and the two channels correspond to the prediction probabilities of the background and cracks, respectively. The 1×1 convolution formula for the grid classification branch is as follows, used to achieve channel dimension compression and mapping: ; In the formula, W is a 1×1 convolution kernel weight matrix, and the parameter dimension is [C]. in C out [, 1, 1], b is the convolution bias, and y is the convolution output feature map. This formula can flexibly adjust the number of channels without changing the feature map spatial size, thereby reducing the computational cost of the model.

3. The method for identifying crack regions in concrete dams based on detection-grid collaboration according to claim 2, characterized in that, The core convolution operation formula for the network is as follows: ; In the formula, y i, j Let k be the (i, j)th pixel value of the convolutional output feature map, and w be the kernel size. m,n Let x be the weight at the (m,n)th position of the convolution kernel. i+m,,j+n is the pixel value at the corresponding position in the input feature map, and b is the convolution bias term; The calculation formula for the channel attention mechanism in the C2PSA module is as follows: ; In the formula, σ is the Sigmoid activation function, MLP is a multilayer perceptron, GlobalAvgPool is a global average pooling operation, x is the input feature map, and Attention(C) is the channel attention weight of the output, which is used to adaptively weight the features of different channels.

4. The method for identifying crack regions in concrete dams based on detection-grid collaboration according to claim 2, characterized in that, The specific bidirectional fusion of the feature pyramid-path aggregation of the neck is as follows: Top-down path: The deep features of the 10th layer output by the backbone network are concatenated with the features of the 6th layer after 1×1 convolution and nearest neighbor upsampling, and then refined by the C3K2 module to obtain the features of the 14th layer; the features of the 14th layer are concatenated with the features of the 4th layer after 1×1 convolution and upsampling, and then refined by the C3K2 module to obtain the features of the 18th layer. Bottom-up path: The features of layer 18 are downsampled by 3×3 convolution and then concatenated with the features of layer 16. The features of layer 21 are then reconstructed by the C3K2 module. The features of layer 21 are downsampled by 3×3 convolution and concatenated with the features of layer 10. The features of layer 24 are then reconstructed by the C3K2 module. The features from layers 18, 21, and 24 are used as inputs to the object detection branch, while the features from layer 21 are used as inputs to the grid classification branch.

5. The method for identifying crack regions in concrete dams based on detection-grid collaboration according to claim 1, characterized in that, The dynamic generation of grid truth labels in step S2 specifically includes: S21. Divide the input image into 40×40 grid units evenly. The size of each grid unit is 1 / 40×1 / 40 in the normalized coordinate system. S22. Read the normalized detection box parameters from the annotation file, including center coordinates, width and height, and category; S23. Traverse each grid cell and calculate its intersection area ratio (IoA) with all detection boxes, which is the intersection area of ​​the grid and the detection box divided by the area of ​​the grid itself. Assign crack or background categories according to the preset label threshold. If the maximum IoA is greater than the label threshold, mark it as crack with a corresponding mapping value of 1. Otherwise, mark it as background with a corresponding mapping value of 0. Generate the full-image grid label matrix. S24. During the data augmentation process, the grid labels are recalculated in real time along with the transformed detection box, automatically maintaining the same displacement, scaling, flipping, and clipping changes as the detection box, without the need for additional alignment operations.

6. The method for identifying crack regions in concrete dams based on detection-grid collaboration according to claim 1, characterized in that, The joint loss function in step S3 is as follows: S31 Total Loss Formula: ; In the formula L det It is to detect loss, L grid For grid classification loss, λ grid Weights for grid classification loss; Detection loss L det The original v8DetectionLoss loss function of YOLO11 is adopted, which includes bounding box regression loss, class loss and distribution focus loss; Grid classification loss L grid A weighted combination of Dice Loss and Focal Loss is used, as shown in the formula: ; The Dice Loss function is expressed as follows: ; In the formula, ε = 1 × 10 -6 w represents the pixel weight, with a foreground pixel weight of 7.0 and a background pixel weight of 0.5; p represents the grid prediction probability value, and t represents the grid's true label. The Focal Loss function is expressed as follows: ; In the formula, α t =0.6, γ =2.0, P t The predicted probability for the correct category; After the network output is parsed and separated, the detection feature list is fed into v8DetectionLoss to calculate the three components of the detection loss and the total detection loss L. det The grid prediction tensor, after being processed by softmax to obtain the positive class probability, is fed into GridDiceFocalLoss along with the grid labels to calculate L. grid The final total loss L is obtained. total .

7. The method for identifying crack regions in concrete dams based on detection-grid collaboration according to claim 1, characterized in that, The MaM collaborative reasoning algorithm in step S4 specifically includes: S41. After processing the prediction results output by the grid classification branch with softmax, the positive class probability is taken to obtain the grid score map gridscores. Grids that are predicted to be positive and whose probability is greater than the confidence threshold are marked as positive grids. S42. For each initial positive grid, calculate the intersection area ratio IoA with all detection boxes. If IoA>0 and the corresponding detection box category is crack, then directly retain it as a crack grid. S43. For the initial positive grid that is not covered by any detection box, extract the corresponding image patch in the original image and calculate the local texture intensity score: ; In the formula, GM is the average gradient magnitude calculated by the Sobel operator; if the score is higher than the texture threshold, it is retained as a cracked mesh, otherwise it is suppressed as background. S44. After two rounds of screening, the remaining crack meshes form the corrected binary mesh diagram.

8. The method for identifying crack regions in concrete dams based on detection-grid collaboration according to claim 1, characterized in that, The connected component merging and quantization output in step S5 specifically include: S51. After normalizing all detection boxes, map them to a 40×40 grid space. Mark all grid cells that intersect with the detection boxes as positive samples. Perform an OR operation with the corrected binary grid map and merge them to obtain the total binary grid map. S52. Use the 8-neighborhood connected component labeling algorithm to perform connected component analysis on the total binary grid graph and assign a unique identifier to each independent region. S53. Post-process the connected component results, filter out connected regions that do not contain any detection boxes, and merge overlapping connected regions to obtain the final crack connected region. S54. For each final crack connected region, calculate its row and column boundaries in the grid space, convert them into pixel coordinates of the input image, generate the bounding box of the corresponding crack region, count the number of cracks and output the location information.

9. The method for identifying crack regions in concrete dams based on detection-grid collaboration according to claim 2, characterized in that, The weight initialization of the dual-task detection model adopts a differentiated strategy: all convolutional layers are initialized with the Kaiming normal distribution to adapt to the SiLU activation function; all BatchNorm2d layer weights are initialized to 1 and biases are initialized to 0; the bias of the last 1×1 convolution layer of the grid classification branch is set differently, with the negative class background bias initialized to 2.0 and the positive class crack bias initialized to -1.0, so that the grid branch output is biased towards the background in the early stage of training to increase the gradient backpropagation of positive samples.

10. The method for identifying crack regions in concrete dams based on detection-grid collaboration according to claim 1, characterized in that, The training process of the dual-task detection model is as follows: The training dataset consists of a mixture of images taken on-site at concrete dams and a publicly available crack database, divided into training, validation, and test sets in a 7:2:1 ratio. All images are labeled only with crack detection boxes. The Adam optimizer was used, with an initial learning rate of 0.

001. Cosine annealing was employed to decay the learning rate, with a momentum of 0.937 and a weight decay of 0.

01. A learning rate warm-up of 10 rounds was set. The training batch size is 15, the total number of training rounds is 150, and the early stop tolerance is set to 70. During training, multiple data augmentation strategies, including color transformation, geometric transformation, mosaic, mixup, and copy-paste, are used to improve the model's generalization ability.