Dam crack detection method and device based on multi-dimensional and multi-scale feature fusion
By introducing a multi-dimensional multi-scale feature fusion module in the YOLOv8 model, the difficulty of identifying traditional dam crack detection methods in complex backgrounds and low light conditions is solved, and efficient and accurate crack detection effect is achieved.
Patent Information
- Application Number
- CN202510009446.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
AI Technical Summary
Traditional dam crack detection methods are inefficient, inaccurate and costly, and it is difficult to accurately identify small target cracks in complex backgrounds and low light conditions.
Using a multi-dimensional multi-scale feature fusion method based on the YOLOv8 model, an improved YOLOv8 model is constructed through the TDMCA attention mechanism, the SPDConv module, the C2f-AKKDBB module and the GP-Detect detection head, and an improved YOLOv8 model is used to identify dam cracks in images taken by the drone.
It significantly improves the accuracy and efficiency of crack detection, and can accurately identify small target cracks under complex backgrounds and low light conditions, meeting the real-time requirements of drone remote sensing image processing.
Smart Images

Figure CN119941668A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of dam health monitoring, and in particular relates to a dam crack detection method and device based on multi-dimensional and multi-scale feature fusion. Background Art
[0002] my country has abundant hydropower resources and is the world's largest hydropower producer. As an important part of my country's energy structure, hydropower not only provides strong driving force for the country's sustainable development, but also plays a vital role in promoting economic growth and social progress. In this context, dams, as the core infrastructure for hydropower development, are directly related to national energy security and economic and social stability in their construction and operation. Dams not only undertake multiple functions such as power generation, irrigation, and flood control, but also provide important guarantees for regional economic development and ecological and environmental protection.
[0003] However, over time, the structure and system of the dam are inevitably eroded by various environmental factors. When hydraulic concrete structures are exposed to the natural environment for a long time, they are affected by factors such as water scouring, temperature changes, freeze-thaw cycles, chemical erosion, etc., which leads to the gradual decline of material properties. In addition, the repeated action of loads and the aging process of materials also gradually weaken the resistance of the dam. These damage and degradation processes may not be easy to detect in the short term, but over time, their cumulative effects will significantly reduce the ability of the dam to resist natural disasters, and may even lead to the degradation of the dam's structural functions under normal operating conditions. If effective maintenance and monitoring measures are not taken in time, in extreme cases, these problems may cause serious safety accidents and lead to catastrophic consequences.
[0004] Various diseases and defects of hydraulic structures usually first appear on the surface of the structure, such as cracks, damage, abrasion, leakage and steel corrosion. Cracks are obvious signals of potential safety hazards in dam structures, because cracks form mechanical discontinuities inside the dam body, destroying the integrity of the dam and significantly reducing its bearing capacity.
[0005] Therefore, timely detection of dam cracks is crucial to improving the early warning function of the dam safety monitoring system. Accurate crack detection data is the core basis for dam safety diagnosis, but how to obtain this data efficiently and accurately is still a difficult problem that needs to be solved urgently.
[0006] Traditional dam crack detection mainly relies on manual detection methods, where inspectors directly observe cracks and other defects on the dam surface through human eyes or small auxiliary equipment. However, this method has many defects: low detection efficiency, incomplete and inaccurate data collection, and is easily affected by the subjective judgment of inspectors. In addition, manual detection is costly, especially when dealing with large-scale projects. These factors make it difficult for traditional detection methods to meet the needs of modern dam safety monitoring. In recent years, vision-based automated crack detection technology has gradually attracted attention. This type of technology collects images of the dam surface through equipment such as drones, and uses computer vision algorithms to identify and analyze cracks, which has the characteristics of non-destructive, high precision and high efficiency. However, in practical applications, how to quickly and accurately identify cracks remains a challenge. The single-stage detection algorithm directly predicts the location and category of the target from the image, has a fast inference speed, and is very suitable for application scenarios with high real-time requirements. Considering that the proportion of small targets in the images taken by drones is high and the background is complex, based on the requirements of detection accuracy and processing speed, the real-time requirements of drone remote sensing image processing can be better met. Summary of the invention
[0007] The technical problem to be solved by the present invention is to address the deficiencies involved in the background technology and provide a UAV crack detection method based on the multi-dimensional and multi-scale feature fusion of the YOLOv8 model, which can effectively solve the problem that the original YOLOv8 detection network is difficult to accurately identify the target due to the high proportion of small targets, large target scale differences, and complex background in the UAV images, thereby improving the detection accuracy and efficiency and meeting the real-time requirements of UAV remote sensing image processing.
[0008] The technical solution adopted by the present invention is:
[0009] In a first aspect, the present invention protects a dam crack detection method based on multi-dimensional and multi-scale feature fusion, the detection method comprising the following steps:
[0010] Step 1: Collect crack images of reinforced concrete dams through drones, and manually annotate the images using annotation software to obtain a data set;
[0011] Step 2: Build an improved YOLOv8 model and use the dataset to train and evaluate the improved YOLOv8 model;
[0012] Step 3: Input the dam crack image to be detected into the trained improved YOLOv8 model for recognition to obtain the crack detection results.
[0013] Preferably, in step 1, the collected dam body crack graphics are first annotated with rectangular frames using labelimg annotation software, and the data set is amplified and then divided into a training set, a validation set, and a test set according to a ratio of 7:2:1.
[0014] More preferably, the improved YOLOv8 model construction process in step 2 is as follows:
[0015] Step 201: Build a YOLOv8 model, which is mainly composed of a backbone network, a neck network and a detection head, and use the YOLOv8 model as a basic model;
[0016] Step 202: The C2f module in the backbone network is replaced with the C2f-AKDBB module, and the TDMCA attention mechanism is added after the 4th, 6th, and 9th layers. Then, the standard convolutions of the 3rd, 5th, 7th, 19th, and 22nd layers are replaced with SPDConv, and the detection heads after the 18th, 21st, and 24th layers are replaced with the GP-Detect detection head to obtain an improved YOLOv8 model.
[0017] In the second aspect, the present invention protects a dam crack detection device based on multi-dimensional and multi-scale feature fusion, which includes a memory deployed with an improved YOLOv8 model and a processor for executing computer programs. The processor is mounted on an intelligent drone, and the intelligent drone client is connected to a workstation on the ground via a WiFi router. During the aerial photography of the intelligent drone, the surface of the dam is photographed, and the video images taken by the camera are directly input into the deployed processor for processing, and cracks on the dam surface are detected by using the improved YOLOv8 model.
[0018] Beneficial effects:
[0019] The present invention adopts multiple innovative modules, including the TDMCA attention mechanism, which extracts image semantic information across dimensions in the X, Y, and Z dimensions, and performs branch splitting and parallel processing of features in stages (such as B-stage, C-stage, etc.). It also includes an SPDConv module, which combines the space-depth layer and the stepless convolution layer to optimize the detection of small targets under low light and complex background conditions. This module significantly improves the detection accuracy under low-resolution conditions through deeper feature extraction and spatial detail retention. It also includes a C2f-AKKDBB module, which is based on the combined design of AKconv and DBB to enhance the feature extraction and expression capabilities of the model. AKconv extracts efficient channel features and optimizes the allocation of computing resources; DBB extracts multi-scale spatial features through a multi-branch structure, and uses structural reparameterization in the inference stage to effectively control the computing cost. It also includes a GP-Detect module, which combines PConv and Group Conv to reduce the model calculation burden and improve the efficiency of the detection head. Through multi-module innovative design, the feature expression and fusion capabilities of YOLOv8 in drone scenarios have been improved, making it more suitable for small target detection in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0021] Figure 1 is a network structure diagram of the present invention;
[0022] Figure 2 is a structural diagram of TDMCA in the present invention;
[0023] Figure 3 It is a DBB structure diagram of the present invention;
[0024] Figure 4 This is a schematic diagram of labelimg annotation data set used in the present invention;
[0025] Figure 5 It is a schematic diagram of the dam crack detection results in the present invention.
[0026] Figure 6 It is the muddy sand irrigation area-tank section-the first section of the retaining wall.
[0027] Figure 7 It is the muddy sand irrigation area-tank section-the second section of the retaining wall.
[0028] Figure 8 It is the muddy sand irrigation area-channel side wall section-the first section of the retaining wall.
[0029] Fig. 9 It is the muddy sand irrigation area-channel side wall section-the first section of the retaining wall. DETAILED DESCRIPTION
[0030] Now, various exemplary embodiments of the present invention are described in detail, and this detailed description should not be considered as a limitation of the present invention, but should be understood as a more detailed description of certain aspects, characteristics and embodiments of the present invention. It should be understood that the terms described in the present invention are only for describing specific embodiments and are not used to limit the present invention.
[0031] In addition, for the numerical range in the present invention, it is understood that each intermediate value between the upper and lower limits of the range is also specifically disclosed. Each smaller range between the intermediate value in any stated value or stated range and any other stated value or intermediate value in the range is also included in the present invention. The upper and lower limits of these smaller ranges can be independently included or excluded in the scope.
[0032] Unless otherwise indicated, all technical and scientific terms used herein have the same meanings as those generally understood by those skilled in the art. Although the present invention describes only preferred methods and materials, any methods and materials similar or equivalent to those described herein may also be used in the implementation or testing of the present invention. All documents mentioned in this specification are incorporated by reference to disclose and describe the methods and / or materials associated with the documents. In the event of a conflict with any incorporated document, the content of this specification shall prevail.
[0033] It will be apparent to those skilled in the art that various modifications and variations may be made to the specific embodiments of the present invention without departing from the scope or spirit of the present invention. Other embodiments derived from the present invention will be apparent to those skilled in the art. The present invention description and examples are exemplary only.
[0034] Example 1
[0035] This embodiment specifically discloses a dam crack detection method based on multi-dimensional and multi-scale feature fusion, which includes the following steps:
[0036] Step 1: Collect crack images of reinforced concrete dams through drones, and manually annotate the images using annotation software to obtain a data set;
[0037] Specifically, drones were used to take pictures of cracks on the surface of the dam or record videos of cracks in real engineering conditions. The captured videos were framed to obtain crack images, and the images were annotated with rectangular frames using labelimg annotation software to obtain a data set. Sample amplification was performed on the data set to significantly enhance the generalization ability of the model and suppress the overfitting phenomenon of the model. The dam crack data set was divided into training set, validation set, and test set in a ratio of 7:2:1 to ensure the training effect and evaluation performance of the model.
[0038] Step 2: Build an improved YOLOv8 model and use the dataset to train and evaluate the improved YOLOv8 model;
[0039] First, construct a three-dimensional multi-channel coordinate attention, Three-Dimensional Multi-Channel Coordinate Attention (TDMCA) (such as Figure 2 As shown in the figure); it can not only capture the feature position encoding information and direction perception information in the X and Y directions, but also effectively extract the channel feature information in the Z direction. By designing without significantly increasing the amount of computation, TDMCA significantly enhances the information interaction capability within the model, especially in the object detection task under complex backgrounds, improving the accuracy and efficiency of feature extraction;
[0040] Secondly, we use SPDConv to replace some standard convolutional layers to capture more features in the downsampling process while reducing the amount of model computation, thereby speeding up training and inference.
[0041] Then, in the C2f module, a more efficient AKDBB_Bottleneck structure is used to optimize the bottleneck. By combining AKConv and DBB re-parameterization and superposition, the computational complexity of the model is effectively reduced, thereby achieving lightweight model design and improving detection performance;
[0042] Finally, the GP-Detect detection head is proposed, which combines PCone and group convolution techniques to optimize the performance of the detection head module. By reducing redundant calculations and memory accesses, this method significantly improves the calculation speed. GP-Detect shows higher accuracy when processing small target detection, effectively reduces the missed detection rate of small targets, and further enhances the application ability of the model in complex scenarios.
[0043] Specifically, the C2f module in the YOLOv8 backbone network is replaced with the C2f-AKDBB module, and the TDMCA attention mechanism is added after the 4th, 6th, and 9th layers. The standard convolutions in the 3rd, 5th, 7th, 19th, and 22nd layers are replaced with SPDConv. The detection heads after the 18th, 21st, and 24th layers are replaced with the GP-Detect detection head, and a multi-dimensional and multi-scale feature fusion UAV crack detection model is obtained, as shown in the figure. Figure 1 shown.
[0044] The advantages of the above model improvements are:
[0045] TDMCA module (cross-dimensional multi-channel attention mechanism): The TDMCA module extracts image semantic information across dimensions in the X, Y, and Z dimensions, and splits and processes features in branches and in parallel in stages (such as B-stage, C-stage, etc.). This module design improves the model's ability to express features and interact with information in different directions, and is suitable for multi-target detection in complex backgrounds.
[0046] SPDConv module (spatial-depth convolution): The SPDConv module combines the spatial-depth layer and the strideless convolution layer to optimize small target detection under low light and complex background conditions. This module significantly improves detection accuracy under low-resolution conditions through deeper feature extraction and spatial detail preservation.
[0047] C2f-AKKDBB module: Based on the combined design of AKconv and DBB (Diverse Branch Block), it enhances the model's feature extraction and expression capabilities. AKconv extracts efficient channel features and optimizes the allocation of computing resources; DBB (such as Figure 3 As shown in the figure, multi-scale spatial features are extracted through a multi-branch structure, and structural re-parameterization is used in the inference stage to effectively control the computational cost.
[0048] GP-Detect module: GP-Detect combines PConv and Group Conv to reduce the model calculation burden and improve the efficiency of the detection head. PConv applies convolution to some channels to keep it lightweight; Group Conv enhances the information interaction between channels and enriches the feature expression.
[0049] The specific process of training and evaluating the improved YOLOv8 model using the dataset is as follows:
[0050] Data preparation: Each image in the dataset needs to have a corresponding annotation file. The annotation file is a .txt file that contains the location (center coordinates, width, and height) of each object and its category.
[0051] Data format and input: Use the yaml configuration file to specify the paths of the training, validation, and test sets and the category names of the cracks.
[0052] Training process: Use the improved YOLOv8 model to train the training set. The model learns the target detection task through the back-propagation algorithm, and adjusts the internal parameters in each iteration to make the model prediction closer to the true value. The key parameters in the training process are: epochs: the number of training rounds, batch: the number of images in each batch, close_mosaic: Mosaic data enhancement (to improve the ability to detect small targets), optimizer: specify the optimizer type (SGD optimizer is used in the experiment), amp: mixed precision training (to speed up the training process and save video memory).
[0053] Verification and adjustment: After each training phase or a certain period (epoch), the model is evaluated using a validation set. The purpose of the validation set is to monitor the performance of the model on unseen data, thereby preventing the model from overfitting to the training data. By calculating the evaluation indicators (mean average precision mAP, accuracy, recall, etc.) on the validation set, hyperparameters such as learning rate and batch size can be adjusted to optimize model performance.
[0054] Testing and final evaluation: Use the test set to conduct a final evaluation of the model to test the generalization ability of the model. The evaluation results are:
[0055] mAP (Mean Average Precision): reflects the average value of model detection accuracy.
[0056] Precision: Indicates the ratio of correctly detected targets to all detection results.
[0057] Recall: It represents the ratio of correctly detected targets to all true targets.
[0058] F1-score: A comprehensive evaluation indicator of precision and recall.
[0059] Model reasoning: After training and evaluation are complete, the trained model can be used for reasoning, that is, detecting objects on new images. The reasoning results include the category, location, and confidence of each object in the image. The model is further analyzed based on the test results. If the model performs poorly in certain categories or specific scenarios, performance can be optimized through further training, data enhancement, or using a different model architecture.
[0060] Step 3: Input the dam crack image to be detected into the trained improved YOLOv8 model for recognition to obtain the crack detection results.
[0061] Example 2
[0062] This embodiment discloses a dam crack detection device based on multi-dimensional and multi-scale feature fusion, which includes a memory deployed with the improved YOLOv8 model constructed in Example 1 and a processor for executing a computer program. The processor is mounted on an intelligent drone, and the intelligent drone client is connected to a workstation on the ground through a WiFi router. During the aerial photography of the intelligent drone, the surface of the dam is photographed, and the video image captured by the camera is directly input into the deployed processor for processing, and the cracks on the dam surface are detected by the improved YOLOv8 model.
[0063] The above-mentioned intelligent drone uses the DJI Mavic 3E drone. The DJI Mavic 3E drone is used as a flight platform, combined with a high-precision RTK positioning system and a 4 / 3-inch CMOS mechanical shutter sensor to achieve centimeter-level precision aerial survey and image data collection. Through automated route planning and real-time orthophotography functions, high-resolution images can be quickly acquired, which is suitable for application scenarios such as geographic information mapping and 3D modeling. In addition, the DJI Mavic 3E integrates a telephoto lens and intelligent obstacle avoidance technology to meet the needs of long-range target detection and complex environment operations, effectively improving data collection accuracy, flight safety and mission execution efficiency.
[0064] The entire shooting process of the above-mentioned intelligent drone can be divided into:
[0065] (1) Information collection: The high-resolution camera carried by the drone takes real-time photos of the dam surface to obtain high-definition images and video data. The precise positioning function of the DJI Mavic 3E drone is used to ensure high accuracy and comprehensiveness of image collection, especially when shooting in complex terrain or difficult areas on the dam surface, to ensure comprehensive data coverage.
[0066] (2) Information processing: The images captured by the camera are transmitted in real time to the high-performance processor deployed on the drone, and image processing is performed by integrating the improved YOLOv8 model. The model can automatically detect and mark cracks on the dam surface, and classify and accurately locate cracks based on image features. The image recognition capability of deep learning improves the accuracy of crack recognition, processes image data in real time, reduces manual intervention, and improves work efficiency.
[0067] (3) Information transmission: The processed image data and crack detection results are transmitted to the ground workstation in real time via the wireless network. The staff can view the monitoring results and the distribution of cracks on the dam surface in real time on the workstation screen. The local area network ensures high speed and low latency of data transmission, ensuring that on-site workers can obtain the latest monitoring data at any time and respond quickly.
[0068] (4) Information storage and recording: The system automatically records detailed information on the surface cracks of the dam in each detection area, including the location, coordinates, confidence value of the cracks, date and time of detection, etc. All these data will be saved in real time in a text file or database. In order to ensure the integrity and security of the data, all historical detection data will be backed up and stored in the cloud server to facilitate subsequent query and data analysis. Through linkage with the monitoring system, the system can realize intelligent identification, precise positioning and real-time monitoring of cracks on the dam surface, providing continuous support for long-term safety management and maintenance.
[0069] Application Cases
[0070] The deployed improved YOLOv8 model was used to conduct an application test on the retaining wall of the large-scale muddy sand irrigation project to verify its actual detection effect. As a key water conservancy structure in the irrigation area, the detection and monitoring of surface cracks on the retaining wall is crucial to the safe operation of the project.
[0071] First, a DJI Mavic 3Enterprise drone was used to collect images of the retaining wall surface. The drone is equipped with a 4K high-definition camera and an RTK high-precision positioning system, which can fully cover the surface structure of the retaining wall during low-altitude flight and accurately record the geographic location information of each image. In actual operation, the drone takes panoramic images and local close-ups of the retaining wall in sequence according to the predetermined flight path to ensure that the details of the cracks are fully captured.
[0072] The captured images are wirelessly transmitted to the drone’s onboard processor in real time, where an improved YOLOv8 crack detection model is deployed. The model analyzes the images in real time during flight, marks the cracks in the images, and outputs confidence information.
[0073] To train and optimize the model, the surface images of the retaining wall were manually annotated using the LabelImg tool to generate a high-quality data set including crack type, location, size, etc. After the annotation, the YOLOv8 model was trained and parameter optimized for multiple rounds based on these data to ensure that the model can maintain high-precision detection performance under different lighting conditions and complex backgrounds.
[0074] During the test, the model detection results are transmitted to the ground workstation in real time through the wireless module of the drone, and the staff can intuitively view the detection information of each crack on the workstation screen. At the same time, the detection data is automatically stored, including key information such as the geographical location of the crack, confidence value, and detection time. These data are saved as text files for subsequent analysis and comparison.
[0075] After actual testing, the improved YOLOv8 model can quickly and accurately detect multiple cracks on the surface of the retaining wall and mark the detailed information of the cracks (such as Figure 6-9 The test results show that the confidence level of crack detection is over 95%, and the recognition capability of small cracks and complex background cracks is significantly improved, which provides effective support for crack monitoring of retaining walls.
[0076] The application of the present invention in the retaining wall of the irrigation section of a large-scale muddy sand irrigation project has fully verified its applicability and reliability in actual water conservancy projects, which helps to improve the efficiency and accuracy of crack detection in water conservancy projects and ensure the long-term safe operation of the project.
[0077] The embodiments described above are only preferred specific implementation modes of the present invention, and the protection scope of the present invention is not limited thereto. Any simple changes or equivalent replacements of the technical solutions that can be obviously obtained by any technician familiar with the field within the technical scope disclosed in the present invention belong to the protection scope of the present invention.
Claims
1. A dam crack detection method based on multi-dimensional and multi-scale feature fusion, characterized in that: The detection method comprises the following steps: Step 1: Collect crack images of reinforced concrete dams through drones, and manually annotate the images using annotation software to obtain a data set; Step 2: Build an improved YOLOv8 model and use the dataset to train and evaluate the improved YOLOv8 model; Step 3: Input the dam crack image to be detected into the trained improved YOLOv8 model for recognition to obtain the crack detection results.
2. A dam crack detection method based on multi-dimensional and multi-scale feature fusion according to claim 1, characterized in that: In step 1, the collected dam crack graphics are first annotated with rectangular frames using labelimg annotation software, and the data set is amplified and then divided into training set, validation set, and test set according to the ratio of 7:2:
1.
3. The dam crack detection method based on multi-dimensional and multi-scale feature fusion according to claim 1 is characterized in that: The improved YOLOv8 model construction process in step 2 is as follows: Step 201: Build a YOLOv8 model, which is mainly composed of a backbone network, a neck network and a detection head, and use the YOLOv8 model as a basic model; Step 202: The C2f module in the backbone network is replaced with the C2f-AKDBB module, and the TDMCA attention mechanism is added after the 4th, 6th, and 9th layers. Then, the standard convolutions of the 3rd, 5th, 7th, 19th, and 22nd layers are replaced with SPDConv, and the detection heads after the 18th, 21st, and 24th layers are replaced with the GP-Detect detection head to obtain an improved YOLOv8 model.
4. A dam crack detection device based on multi-dimensional and multi-scale feature fusion, characterized in that: The device includes a memory deployed with the improved YOLOv8 model constructed according to claim 3 and a processor for executing a computer program. The processor is mounted on an intelligent drone. The intelligent drone client is connected to a workstation on the ground via a WiFi router. During the aerial photography process of the intelligent drone, the surface of the dam is photographed, and the video images taken by the camera are directly input into the deployed processor for processing, and cracks on the dam surface are detected by the improved YOLOv8 model.
Citation Information
Cited By
Building surface defect detection method and system based on improved YOLOv8
CN121481958A