Crack detection model training method, crack detection method and device
Patent Information
- Application Number
- CN202511507932.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-21
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-10-21
AI Technical Summary
[0003]然而,传统的基于深度学习的机场跑道裂缝检测方法的检测精度与检测速度不高,会导致检测效率较低,同时会导致最终检测到的裂缝识别结果的准确性较低
[0010]本申请实施例提供的裂缝检测模型训练方法、裂缝检测方法及装置,通过在裂缝检测模型中采用LSAF模块,能够采用低秩自注意力机制捕捉长距离依赖,采用标准卷积提取局部特征,二者结合让该LSAF模块能同时学习到特征的长距离关系与局部细节,丰富特征表达,并且,低秩注意力机制通过降维计算,能够减少自注意力计算量和内存占用,降低了计算复杂度,以有效提高机场跑道的裂缝识别结果的检测效率,此外,在裂缝检测模型中添加DC-AFM,能够通过使用不同膨胀率的空洞卷积,有效获取图像多尺度特征,满足对不同大小目标的检测需求,使裂缝检测模型在复杂场景中能兼顾目标的细节特征和整体结构特征,提升特征表示的丰富性,同时,从通道和空间这两个维度筛选和增强特征,抑制冗余信息,让裂缝检测模型更精准定位和识别目标,进一步提升特征针对性和有效性,以有效提高裂缝识别结果的检测准确性。也就是说,整个方法采用裂缝检测模型,能够快速且准确地检测出机场跑道的裂缝识别结果,即能够有效提高对机场跑道裂缝识别精度和识别速度。
Smart Images

Figure CN121353230B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image data processing technology, and in particular to a crack detection model training method, crack detection method and apparatus. Background Technology
[0002] Crack detection on airport runways is a crucial step in the safety assessment of airport pavement structures. This step enables the timely detection of potential hazards on runways, ensuring their safe operation. Crack detection allows electronic equipment (such as pavement health monitoring systems) to monitor the structural condition of the runways in real time, accurately identifying various types of cracks, such as stress cracks, temperature cracks, and shrinkage cracks. This provides important information for runway maintenance and reinforcement decisions. However, due to the small size of cracks and limitations imposed by the complex testing environment and equipment performance, detecting small-target cracks is often challenging. Therefore, the ability to quickly and accurately detect small-target runway cracks is of great significance for ensuring safe airport operation and scientifically formulating runway maintenance strategies, possessing high practical value and promising application prospects.
[0003] However, traditional deep learning-based methods for detecting cracks in airport runways suffer from low accuracy and speed, resulting in low detection efficiency and consequently, low accuracy in the final crack identification results. Therefore, how to quickly and accurately detect cracks in airport runways has become a pressing issue. Summary of the Invention
[0004] This application provides a crack detection model training method, a crack detection method, and an apparatus to effectively improve the detection efficiency and accuracy of crack identification results for airport runways.
[0005] In a first aspect, embodiments of this application provide a method for training a crack detection model, wherein the crack detection model includes: a multi-scale image determination module, a first image processing module, a first low-rank self-attention fusion (LSAF) module, a second image processing module, a dilated convolution-attention fusion (DC-AFM) module, and a third image processing module; the method includes: using the multi-scale image determination module to determine a first image sample, a second image sample, and a third image sample based on initial test image samples of an airport runway sample; using the first image processing module to improve the resolution of the first image sample and then stitch it with the second image sample to obtain a first stitched image sample; and using the first LSAF module to capture the long-distance dependencies between features in the first stitched image sample to obtain a first feature image sample. The second image processing module is used to improve the resolution of the first feature image sample and then stitch it with the third test image sample to obtain the second stitched image sample. DC-AFM is then used to extract, filter, and enhance the detailed features of the second stitched image sample to obtain the second feature image sample. The third image processing module is used to determine the first target feature image sample corresponding to the first test image sample, the second target feature image sample of the second test image sample, and the third target feature image sample of the third test image sample based on the first test image sample, the first feature image sample, and the second feature image sample. Based on the first target feature image sample, the second target feature image sample, the third target feature image sample, and the actual crack feature image sample, the parameters of the crack detection model are updated to obtain the trained crack detection model.
[0006] Secondly, embodiments of this application provide a crack detection method, comprising: acquiring a test image of an airport runway; inputting the test image into a crack detection model to obtain a crack identification result output by the crack detection model, the crack identification result including: crack bounding box and crack classification result; wherein, the crack detection model is obtained by training based on the crack detection model training method of the first aspect.
[0007] Thirdly, embodiments of this application also provide a crack detection model training device, comprising: The multi-scale image determination module is used to determine the first image sample, the second image sample, and the third image sample based on the initial image sample to be tested from the airport runway sample. The first image processing module is used to improve the resolution of the first image sample to be tested and then stitch it with the second image sample to be tested to obtain the first stitched image sample. The first LSAF module is used to capture the long-distance dependencies between features in the first stitched image sample to obtain the first feature image sample; The second image processing module is used to improve the resolution of the first feature image sample and then stitch it with the third image sample to be tested to obtain the second stitched image sample. DC-AFM is used to extract, filter, and enhance the detailed features of the second stitched image sample to obtain the second feature image sample. The third image processing module is used to determine, based on the first image sample to be tested, the first feature image sample, and the second feature image sample, the first target feature image sample corresponding to the first image sample to be tested, the second target feature image sample of the second image sample to be tested, and the third target feature image sample of the third image sample to be tested; The parameter update module is used to update the parameters of the crack detection model based on the first target feature image sample, the second target feature image sample, the third target feature image sample and the actual crack feature image sample, so as to obtain the trained crack detection model.
[0008] Fourthly, embodiments of this application also provide a crack detection device, comprising: The image acquisition module is used to acquire images of the airport runway to be tested. The crack recognition module is used to input the image to be tested into the crack detection model and obtain the crack recognition result output by the crack detection model. The crack recognition result includes crack bounding boxes and crack classification results. The crack detection model is obtained by training based on the crack detection model training method described in the first aspect.
[0009] Fifthly, embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements the crack detection model training method as described in the first aspect or the crack detection method as described in the second aspect.
[0010] The crack detection model training method, crack detection method, and apparatus provided in this application employ an LSAF module in the crack detection model. This module uses a low-rank self-attention mechanism to capture long-distance dependencies and standard convolution to extract local features. The combination of these two methods allows the LSAF module to simultaneously learn the long-distance relationships and local details of features, enriching feature representation. Furthermore, the low-rank attention mechanism reduces the computational cost and memory usage of self-attention through dimensionality reduction, thereby lowering computational complexity and effectively improving the detection efficiency of crack identification results for airport runways. In addition, adding DC-AFM to the crack detection model allows for the effective acquisition of multi-scale image features by using dilated convolutions with different dilation rates. This meets the detection requirements for targets of different sizes, enabling the crack detection model to consider both the detailed features and overall structural features of targets in complex scenes, thus enhancing the richness of feature representation. Simultaneously, by filtering and enhancing features from both channel and spatial dimensions and suppressing redundant information, the crack detection model can more accurately locate and identify targets, further improving feature specificity and effectiveness, and ultimately enhancing the detection accuracy of crack identification results. In other words, the entire method uses a crack detection model, which can quickly and accurately detect the crack identification results of airport runways, thus effectively improving the accuracy and speed of airport runway crack identification. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a schematic diagram of the crack detection model provided in the embodiments of this application; Figure 2 This is one of the flowcharts illustrating the crack detection model training method provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of the multi-scale image determination module provided in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of the first image processing module provided in the embodiments of this application; Figure 5a This is a schematic diagram of the structure of the first LSAF module provided in the embodiments of this application; Figure 5b This is a schematic diagram of the internal structure of a low-rank self-attention branch provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of the second image processing module provided in the embodiments of this application; Figure 7 This is a schematic diagram of the structure of the DC-AFM provided in the embodiments of this application; Figure 8 This is a schematic diagram of the structure of the third image processing module provided in the embodiments of this application; Figure 9 This is a schematic diagram of the structural connection of the crack detection model provided in the embodiments of this application; Figure 10 This is the second flowchart illustrating the crack detection model training method provided in the embodiments of this application; Figure 11 This is a schematic flowchart of the crack detection method provided in the embodiments of this application; Figure 12a This is one of the schematic diagrams of the target feature image provided in the embodiments of this application; Figure 12b This is a second schematic diagram of the target feature image provided in the embodiments of this application; Figure 13 This is a schematic diagram of the structure of the crack detection model training device provided in the embodiments of this application; Figure 14 This is a schematic diagram of the crack detection device provided in the embodiments of this application; Detailed Implementation To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0013] To better understand the embodiments of this application, the background technology will first be described in detail: In existing technologies, traditional deep learning-based object detection methods achieve more accurate, robust, and efficient object detection through multi-layered network structures and extensive data training. Based on whether candidate bounding boxes need to be generated first, deep learning-based object detection methods can be broadly categorized into two-stage methods and one-stage methods. The former generates candidate regions before feature extraction and classification, while the latter performs feature extraction, location regression, and category classification in a single step.
[0014] It's worth noting that typical two-stage methods, such as Faster Region-based Convolutional Neural Network (Faster R-CNN), introduce the concept of Spatial Pyramid Pooling (SPP). This involves pooling regions of interest (RoIs) onto fixed-size feature maps, eliminating unnecessary computational steps and accelerating feature extraction. Simultaneously, it abandons selective search and employs Region Proposal Networks (RPNs) to automatically generate candidate regions, effectively addressing the computational complexity issue in multi-object detection. However, R-CNN-like object detection methods suffer from slow speed and high memory consumption.
[0015] In recent years, with the continuous improvement of deep learning algorithms, deep learning-based object detection methods have made some progress in the field of object detection. However, traditional deep learning-based airport runway crack detection methods often rely on the size, shape, and texture features of defects in airport runway images to detect cracks. But because the size, shape, and texture features of defects vary, and the crack detection process is relatively complex, this leads to low detection efficiency and low accuracy of the final crack identification results. Furthermore, because airport runway cracks are small, the detection accuracy and speed of the aforementioned airport runway crack detection methods are not high. Therefore, there will be some missed detections during the crack detection process, which also leads to low accuracy of the final crack identification results and low overall detection efficiency. Therefore, how to quickly and accurately detect airport runway cracks has become an urgent problem to be solved.
[0016] This application provides a crack detection model training method, crack detection method, and apparatus. By employing an LSAF module in the crack detection model, a low-rank self-attention mechanism can be used to capture long-distance dependencies, and standard convolution can be used to extract local features. The combination of these two allows the LSAF module to simultaneously learn the long-distance relationships and local details of features, enriching feature representation. Furthermore, the low-rank attention mechanism reduces the computational cost and memory usage of self-attention through dimensionality reduction, thereby lowering computational complexity and effectively improving the detection efficiency of crack recognition results for airport runways. In addition, adding DC-AFM to the crack detection model can effectively acquire multi-scale features of the image by using dilated convolutions with different dilation rates, meeting the detection needs of targets of different sizes. This allows the crack detection model to take into account both the detailed features and overall structural features of the target in complex scenes, improving the richness of feature representation. At the same time, features are filtered and enhanced from the channel and spatial dimensions to suppress redundant information, allowing the crack detection model to more accurately locate and identify targets, further improving feature targeting and effectiveness, and thus effectively improving the detection accuracy of crack recognition results. In other words, the entire method uses a crack detection model, which can quickly and accurately detect the crack identification results of airport runways, thus effectively improving the accuracy and speed of airport runway crack identification.
[0017] It should be noted that the execution subject of the crack detection model training method provided in this application embodiment can be a crack detection model training device or an electronic device, and no specific limitation is made here.
[0018] Optionally, the aforementioned electronic devices may include computers, mobile terminals, and wearable devices, etc.
[0019] The following uses an electronic device as an example to illustrate the crack detection model training method provided in this application embodiment: First, this embodiment of the application uses a Cross-Stage Partial Darknet (CSPDarknet) as the backbone network. In the residual branch of the C3 module in the neck network of the backbone network, a low-rank self-attention mechanism is introduced to replace some standard convolution operations. That is, a low-rank self-attention fusion (LSAF) module is used in the neck network. A dilated convolution-attention fusion module (DC-AFM) is also added to the neck network. In addition, the intersection over union (IoU) loss function in the detection head of the backbone network is replaced with a position-enhanced complete IoU (PECIoU) loss function to construct a multi-scale low-rank fusion attention network, that is, to construct a crack detection model.
[0020] The LSAF module is used to capture long-distance dependencies between features in an image, thereby enhancing the ability to extract local details and non-global context features.
[0021] DC-AFM is used to expand the receptive field by employing dilated convolutions with different dilation rates, while combining the advantages of channel attention and spatial attention, significantly improving the ability to extract global features and fuse multi-scale features.
[0022] It should be noted that the original IoU loss function, which linearly scales the aspect ratio of the predicted bounding box with that of the actual bounding box, can lead to relatively large fluctuations in the convergence process, thus affecting the network optimization of the crack detection model. Therefore, the PECIoU loss function can be introduced to comprehensively measure the similarity between the predicted and actual bounding boxes, making the crack detection model converge faster and the regression results more accurate, thereby effectively improving the performance of the target detection task.
[0023] Secondly, before training the crack detection model, electronic devices can first collect a large amount of airport runway crack data (i.e., initial test image samples) using image acquisition devices, and construct an airport runway crack dataset. Then, sample images are selected from the airport runway crack data for airport runway crack data annotation, splitting into image dataset files and corresponding image label files, resulting in an annotated airport runway crack dataset. This annotated airport runway crack dataset is then classified, including the airport runway crack classification results. Furthermore, this annotated airport runway crack dataset is divided into training, testing, and validation sets, and data augmentation is performed. When training the crack detection model, the training and validation sets are input into the constructed crack detection model to improve its performance and generalization ability. During training, the training set is used to optimize the parameters and learn features of the crack detection model, while the evaluation metrics of the validation set are used to adjust the hyperparameters of the crack detection model to improve its prediction performance on unknown image data. Through multiple rounds of iterative training and validation, the crack detection model is optimized. When using a crack detection model, the test set is input into the trained crack detection model to obtain the crack identification results output by the crack detection model.
[0024] Specifically, in the process of constructing an labeled airport runway crack dataset, sample images are selected from a large number of initial test image samples and airport runway crack labels are created using open-source image annotation tools (such as LabelMe). The created labels are then divided into training set, validation set and test set according to a set ratio (such as 7:2:1), and a composite data augmentation strategy is adopted, namely, using the mosaic method to augment the data, in order to expand the dataset and improve the robustness of the crack detection model.
[0025] Specifically, the data augmentation process is as follows: four images are randomly selected using the mosaic method and spliced together through random scaling, random cropping, and random arrangement to generate new training samples, thereby improving the robustness and generalization ability of the crack detection model.
[0026] For example, the mosaic method is used to stitch together four randomly selected 400×400 pixel images, and the blank areas are filled with average grayscale to generate an enhanced 1280×1280 pixel image. At the same time, a variety of enhancement techniques are introduced (such as MixUp, CutMix, and other hybrid enhancement methods), and combined with enhancement techniques such as color space transformation, geometric deformation, and random occlusion to construct a multi-layered data enhancement system.
[0027] It should be noted that the data augmentation process uses the mosaic method to randomly select four images and stitch them together through random scaling, random cropping, and random arrangement to generate new training samples, enriching the dataset and improving the robustness of the crack detection model. In the last ten rounds of training, mosaic data augmentation is turned off, allowing the crack detection model to focus more on the details of real images, effectively improving the performance of the crack detection model in airport runway crack identification and generalization ability.
[0028] Optionally, the aforementioned image acquisition device can also be referred to as an image acquisition module, used to provide high-quality, highly stable raw images and training data for electronic equipment (such as an airport runway crack detection system). Specifically, the image acquisition module can acquire test image samples of airport runway samples, including at least an initial test image sample and multi-scale image samples (i.e., a first test image sample, a second test image sample, and a third test image sample), providing high-quality input data for subsequent crack detection model training and crack identification detection.
[0029] Optionally, the image acquisition module consists of an optical imaging module, an auxiliary sensing module, and a data interface module. These modules work together to ensure the accuracy of the image data and its environmental adaptability.
[0030] The aforementioned optical imaging module employs a global shutter industrial camera with a resolution of no less than 3840×2160 pixels, a frame rate of ≥30fps, and a dynamic range of ≥120dB, in order to adapt to the complex lighting conditions of the airport runway sample.
[0031] The aforementioned auxiliary sensing module is rigidly connected to the optical imaging module to ensure spatiotemporal data synchronization with a synchronization error of ≤1ms.
[0032] The aforementioned data interface module supports Gigabit Ethernet Vision (GV) and Camera Link (CL) protocols, with a transmission bandwidth of ≥1Gbps. Mechanical locking and electromagnetic shielding are achieved through an M12 waterproof aviation connector. This data interface module integrates a hardware encoder, supports H.265 encoding compression with a compression ratio of ≥50:1, reducing data transmission load. Furthermore, this data interface module provides a Global Positioning System (GPS) / BeiDou dual-mode positioning interface, recording the acquisition location information for each frame of image with a positioning accuracy of ≤0.5m.
[0033] Understandably, the workflow of the aforementioned image acquisition module is as follows: When the airport runway crack detection system is activated, the optical imaging module automatically adjusts the focal length and exposure time according to preset parameters to acquire raw image data; simultaneously, the auxiliary sensing module measures distance and attitude data in real time and embeds image metadata; then, the data interface module packages the acquired raw image data and auxiliary data and transmits them to the preprocessing module in the electronic device via Gigabit Ethernet or 5G wireless link. This image acquisition module can operate stably in environments ranging from -20℃ to 60℃, with a protection level of IP67, meeting the requirements of harsh airport runway conditions.
[0034] The entire process utilizes multi-sensor fusion and high-precision synchronization design to ensure that the acquired raw image data meets the input requirements of the preprocessing module, providing a reliable data foundation for subsequent crack feature extraction and model training.
[0035] Optionally, such as Figure 1 The diagram shown is a structural schematic of the crack detection model provided in an embodiment of this application. Figure 1 The crack detection model may include: a multi-scale image determination module 101, a first image processing module 102, a first low-rank self-attention fusion LSAF module 103, a second image processing module 104, a dilated convolution-attention fusion module DC-AFM 105, and a third image processing module 106.
[0036] Combination Figure 1 As can be seen, the aforementioned preprocessing module is used to perform real-time classification processing on the raw image data transmitted by the image acquisition module, generating feature data (i.e., the first image sample to be tested, the second image sample to be tested, and the third image sample to be tested) that can be used for crack detection training. In other words, this preprocessing module is used to process the raw image data in real time to generate multi-scale feature data that meets the input requirements of the crack detection model.
[0037] The preprocessing module adopts an architecture that combines programmable logic devices with a dedicated hardware acceleration engine to realize multi-scale feature data generation, noise reduction filtering and feature pre-extraction functions, providing standardized input for the core computing module.
[0038] It should be noted that the preprocessing module is built on a multi-core processor system platform, comprising a Programmable Logic (PL) section and a Processing System (PS) section. The PL section accelerates the image processing pipeline, while the PS section runs the operating system for task scheduling and data management. The PL section is configured with... Figure 1The diagram shows three main functional units: a multi-scale image generation module 101, a noise reduction and filtering module, and a feature pre-extraction module. The multi-scale image generation module 101 receives raw image data from the image acquisition module and generates three image samples at different scales—a first test image sample, a second test image sample, and a third test image sample—through a bilinear interpolation downsampling unit. The multi-scale image generation module 101 employs a pipelined architecture design, including an input buffer (such as First In First Out, FIFO), a line buffer, and an interpolation calculation unit, supporting real-time processing of images up to 4K resolution.
[0039] The aforementioned preprocessing module is also connected to the core computing module via a high-speed AXI-Stream interface, achieving a data transmission bandwidth of 20Gbps. Internally, the preprocessing module employs a double-buffering mechanism to ensure the continuity of data processing and transmission. In terms of power management, the preprocessing module supports Dynamic Voltage and Frequency Scaling (DVFS), automatically adjusting the operating frequency and voltage according to the processing load, with a typical operating power consumption of 15W. Regarding environmental adaptability, the preprocessing module meets the operating temperature range of -40℃ to 85℃ and has an IP50 protection rating, making it suitable for the harsh environment of airport runway inspection. Through a hardware-based parallel processing architecture, the preprocessing module controls the multi-scale image sample generation latency to within 5ms, providing highly efficient preprocessing capabilities for airport runway crack detection systems.
[0040] The aforementioned core computation module, also known as the crack detection model training module, is used to train the crack detection model using the multi-scale feature data obtained from the prediction ignoring module. Specifically, this core computation module implements the deep feature extraction and multi-scale feature fusion functions in the crack detection model, and specifically executes the computational tasks of the LSAF and DC-AFM modules described in this paper. This core computation module uses a dedicated hardware architecture to achieve long-distance dependency capture between features, multi-scale feature extraction, and feature selection enhancement functions.
[0041] The core computing module is designed based on an Application-Specific Integrated Circuit (ASIC) and integrates two core processing modules: the LSAF computing engine and the DC-AFM computing engine. The LSAF computing engine includes a first convolution module, a low-rank self-attention branch, a standard residual component branch, a first concatenation module, and a dimensionality-upgrading module. The low-rank self-attention branch is equipped with matrix factorization units and self-attention computing units, supporting dimensionality reduction of input features and long-distance dependency modeling. The standard residual component branch contains a 3×3 convolution accelerator and residual connection units for extracting local features and maintaining gradient propagation. The outputs of the two branches are fused by the feature concatenation module to achieve channel-level fusion, and then dimensionality-upgraded by a 1×1 convolution to output the final features. This LSAF computing engine uses a systolic array architecture to accelerate matrix multiplication. Each computing unit contains 128 parallel processing units, supporting single-cycle computation of an 8×8 matrix with a computational precision of 16-bit floating-point numbers.
[0042] The DC-AFM computation engine consists of a multi-scale feature extraction module and a dual-attention fusion module. The multi-scale feature extraction module is configured with three sets of parallel dilated convolution computation units, corresponding to 3×3 dilated convolution operations with dilation rates of 1, 3, and 5, respectively. Each computation unit contains 64 multiply-accumulate units (MACs), supporting the computation of 8×8 feature blocks in a single cycle. The dual-attention fusion module includes a channel attention submodule and a spatial attention submodule. The former generates channel-dimensional attention weights through global average pooling layers and fully connected layers; the latter generates spatial-dimensional attention weights through 7×7 convolutional layers. The outputs of both submodules are used for feature enhancement through element-wise multiplication, and finally, the feature dimensions are adjusted by 1×1 convolution. This DC-AFM computation engine adopts a dataflow architecture design, supporting pipelined processing of feature maps, and can process up to four sets of 256×256 feature maps simultaneously.
[0043] The aforementioned core computing module is equipped with 4MB of on-chip shared cache and 2GB of Low Power Double Data Rate 4 (LPDDR4) synchronous dynamic external memory for storing intermediate feature data and model parameters. It connects to the host system via a PCIe 4.0×16 interface, with a theoretical transmission bandwidth of 32GB / s. Internally, the core computing module employs a hierarchical storage architecture, including a three-level structure of register files, shared memory, and global memory, reducing data access latency through an intelligent data prefetching mechanism. In terms of power management, the core computing module supports dynamic voltage and frequency adjustment and module-level clock gating technology, with a typical operating power consumption of 45W and a peak performance of 8 TOPS (INT8). Regarding environmental adaptability, the core computing module meets the operating temperature range of -40℃ to 85℃ and has a shock resistance of 5Grms, making it suitable for automotive or airborne vibration environments. Through a hardware-based parallel processing architecture and optimized data path design, the core computing module controls the processing latency of the LSAF module to within 6.8ms and the processing latency of the DC-AFM module to within 8.1ms, providing efficient deep feature computing capabilities for the airport runway crack detection system.
[0044] Furthermore, this application embodiment also relates to an edge computing module, used to deploy the trained crack detection model to achieve real-time crack detection of airport runway images. Specifically, this edge computing module is used to deploy the trained crack detection model to achieve real-time crack detection of airport runways. This edge computing module deploys the crack detection model trained by the aforementioned core computing module on an edge computing device to complete real-time processing and crack identification of the image under test.
[0045] The edge computing module consists of a model deployment unit, an image processing unit, and a result output unit. The model deployment unit loads the trained crack detection model parameters, including... Figure 1 The diagram shows the complete network structure and parameters of the multi-scale image determination module 101, the first image processing module 102, the first LSAF module 103, the second image processing module 104, the DC-AFM module 105, and the third image processing module 106. This edge computing module employs model quantization technology to reduce model size and computational resource consumption while ensuring detection accuracy. The quantization process strictly maintains the computational characteristics of each module as described in the patent.
[0046] The image processing module implements the aforementioned crack detection process, including an image acquisition interface, a preprocessing accelerator, and a core computing accelerator. The image acquisition interface supports GigE Vision and Universal Serial Bus 3.0 (USB 3.0) protocols, and can accept input from cameras with a maximum resolution of 4K. The preprocessing accelerator performs image format conversion, resizing, and normalization, outputting tensor data that meets the model's input requirements. The core computing accelerator employs an optimized implementation of the LSAF and DC-AFM modules described in the paper, accelerating feature extraction and fusion calculations through a dedicated instruction set, with a single image processing latency controlled within 50ms. This edge computing module is equipped with 4GB of dedicated video memory, supporting simultaneous processing of multiple frames of image data. The result output unit generates and outputs crack recognition results, including crack bounding boxes and crack classification results.
[0047] This edge computing module also features a ruggedized design, operating in temperatures ranging from -40°C to 70°C, with an IP65 protection rating, meeting the requirements for use in harsh outdoor environments. The module's power consumption does not exceed 30W, and it supports Power over Ethernet (PoE) and a wide DC 12-24V input voltage, facilitating deployment on vehicle-mounted and drone platforms. Hardware-level security encryption mechanisms ensure the security of model parameters and detection data. This edge computing module strictly adheres to the technical solutions described in the embodiments of this application, efficiently implementing crack detection functions at the edge, providing real-time and reliable detection support for airport runway maintenance.
[0048] Combination Figure 1 ,like Figure 2 The diagram shown is a flowchart illustrating the crack detection model training method provided in an embodiment of this application. Figure 2 The method includes the following steps 201-205.
[0049] Step 201: Using a multi-scale image determination module, determine the first image sample, the second image sample, and the third image sample based on the initial image sample to be tested of the airport runway sample.
[0050] Among them, the aforementioned airport runway sample is a core component of the airport, which is a special long strip of artificially laid or naturally formed ground facility used for aircraft take-off, landing and taxiing.
[0051] The aforementioned initial test image samples refer to the raw image data to be detected. This data may contain airport runways or other irrelevant content. It consists of raw images obtained from actual scenes without any preprocessing or feature extraction. These initial test image samples are part of the training set.
[0052] The three test image samples mentioned above—the first test image sample, the second test image sample, and the third test image sample—are the results obtained after processing the initial test image sample at different scales. In other words, the scales of these three test image samples are different.
[0053] The following section elaborates on how the electronic equipment uses a multi-scale image determination module to determine the first, second, and third image samples based on the initial image samples from the airport runway sample: Optionally, such as Figure 3 The diagram shown is a structural schematic of the multi-scale image determination module provided in an embodiment of this application. Figure 3 In this process, the multi-scale image determination module 101 may include: a first-scale image determination module, a second-scale image determination module, and a third-scale image determination module. The electronic device employs the multi-scale image determination module to determine the first, second, and third image samples to be tested based on the initial image samples to be tested from the airport runway sample. This includes: the electronic device employing the first-scale image determination module to determine the third image sample to be tested based on the initial image samples to be tested from the airport runway sample; the electronic device employing the second-scale image determination module to determine the second image sample to be tested based on the third image sample to be tested; and the electronic device employing the third-scale image determination module to determine the first image sample to be tested based on the second image sample to be tested.
[0054] In this embodiment, the first-scale image determination module may include: a first sub-convolution (Conv) module, a second sub-convolution module, a first feature extraction (C3k2) module, a third sub-convolution module, and a second feature extraction (C3k2) module. The electronic device uses the first and second sub-convolution modules to perform two consecutive local feature extractions on the initial test image sample to obtain a first sub-test image sample; uses the first feature extraction module to perform feature enhancement on the first sub-test image sample to obtain a second sub-test image sample; uses the third sub-convolution module to perform local feature extraction on the second sub-test image sample to obtain a third sub-test image sample; and finally, uses the second feature extraction module to perform feature enhancement on the third sub-test image sample to obtain a third test image sample.
[0055] The second-scale image determination module may include a fourth sub-convolution module and a third feature extraction (C3k2) module. The electronic device uses the fourth sub-convolution module to extract local features from the third test image sample to obtain a fourth sub-test image sample; then, the third feature extraction module is used to enhance the features of the fourth sub-test image sample to obtain the second test image sample.
[0056] The third-scale image determination module may include: a fifth sub-convolution module, a fourth feature extraction (C3k2) module, a cross-convolution positional self-attention (C2PSA) module, and a spatial pyramid feature fusion (SPFF) module. The electronic device uses the fifth sub-convolution module to extract local features from the second test image sample, obtaining the fifth sub-test image sample; the fourth feature extraction module enhances the features of the fifth sub-test image sample, obtaining the sixth sub-test image sample; the C2PSA module then denoises the sixth sub-test image sample, obtaining the seventh sub-test image sample; finally, the SPFF module improves the distribution and separability of features in the seventh sub-test image sample, obtaining the first test image sample.
[0057] Step 202: Using the first image processing module, the resolution of the first image sample to be tested is improved and then stitched with the second image sample to be tested to obtain the first stitched image sample; and using the first LSAF module, the long-distance dependence between features in the first stitched image sample is captured to obtain the first feature image sample.
[0058] In step 202, the electronic device uses a first image processing module to improve the resolution of the first image sample to be tested, making the detailed features in the first image sample to be tested richer and clearer. Then, the first image sample to be tested with the improved resolution is stitched together with the second image sample to be tested to obtain the first stitched image sample. Since the first LSAF module is an improved module based on the C3 module for feature processing, it introduces a low-rank attention mechanism to replace some standard convolution operations. It can be understood as an improved module that combines a low-rank self-attention mechanism with a standard convolution structure. It aims to reduce the computational complexity of the self-attention mechanism and improve computational efficiency by using low-rank matrix factorization technology. At the same time, it combines the local feature extraction capability of standard convolution with the long-distance dependency capture capability of self-attention to enhance the richness and effectiveness of feature representation, thereby improving the performance of the crack detection model in tasks such as target detection. In addition, it can achieve the best balance between detection speed and accuracy. It can effectively improve the overall performance of the crack detection model by reducing information loss and enhancing global interactive representation. Therefore, the electronic device can use the first LSAF module to capture the long-distance dependencies between features in the first stitched image sample and quickly obtain the first feature image sample.
[0059] The following section details how the electronic device uses a first image processing module to improve the resolution of the first image sample under test and then stitches it with the second image sample under test to obtain the first stitched image sample: Optionally, such as Figure 4 The diagram shown is a structural schematic of the first image processing module provided in an embodiment of this application. Figure 4 In this process, the first image processing module 102 may include a first upsampling module and a first image stitching module. The electronic device uses the first image processing module to improve the resolution of the first image sample to be tested and then stitches it with the second image sample to be tested to obtain a first stitched image sample. This process may include: the electronic device using the first upsampling module to improve the resolution of the first image sample to be tested to obtain a first resolution image sample; and using the first image stitching module to stitch the first resolution image sample and the second image sample to be tested to obtain a first stitched image sample.
[0060] In the process of determining the first stitched image sample, the electronic device can first use a first upsampling module to improve the resolution of the first image sample to be tested, so that the detailed features in the first image sample to be tested are richer and clearer, and a first resolution image sample is obtained. Then, a first image stitching module is used to stitch the first resolution image sample and the second image sample to be tested to obtain more comprehensive feature information and obtain a first stitched image sample with rich information and high accuracy.
[0061] The following section elaborates on how the electronic device uses the first LSAF module to capture the long-distance dependencies between features in the first stitched image sample, thus obtaining the first feature image sample: In some embodiments, such as Figure 5a The diagram shown is a structural schematic of the first LSAF module provided in an embodiment of this application. Figure 5aIn this process, the first LSAF module 103 may include: a first convolution module, a low-rank self-attention branch, a standard residual component branch, a first stitching module, and a dimensionality-upgrading module. The electronic device employs the first LSAF module to capture long-distance dependencies between features in the first stitched image sample, obtaining a first feature image sample. This may include: the electronic device using the first convolution module to reduce the dimensionality of the first stitched image sample to obtain a first image sample, and then splitting the first image sample to obtain a first sub-image sample and a second sub-image sample; the electronic device using the low-rank self-attention branch to reduce the dimensionality of the first sub-image sample to obtain a third sub-image sample; and then applying a preset dimensionality-downgrading matrix to the third sub-image sample. The system performs mapping to obtain the fourth sub-image sample; captures the long-distance dependencies between features in the fourth sub-image sample to obtain the fifth sub-image sample; and increases the dimensionality of the fifth sub-image sample to obtain the sixth sub-image sample; the electronic device uses standard residual component branching to reduce the dimensionality of the second sub-image sample to obtain the seventh sub-image sample; performs local feature extraction on the seventh sub-image sample to obtain the eighth sub-image sample; and increases the dimensionality of the eighth sub-image sample to obtain the ninth sub-image sample; the electronic device uses a first stitching module to stitch the sixth and ninth sub-image samples, and uses a dimensionality-increasing module to increase the dimensionality of the stitched sub-image sample to obtain the first feature image sample.
[0062] Among them, the aforementioned first stitched image sample , Represents a three-dimensional tensor; Indicates the first stitched image sample The number of channels, i.e., the length; Indicates the first stitched image sample The number of rows, i.e., the height; Indicates the first stitched image sample The number of columns, i.e., the width.
[0063] In determining the first feature image sample, the electronic device first uses a first convolutional module with a 1×1 kernel to process the first stitched image sample. Dimensionality reduction can be performed using the first dimensionality reduction formula. The first image sample was obtained. ,in, This represents the first convolutional module with a 1×1 kernel, i.e., a convolution operation with a 1×1 kernel. This effectively reduces the computational load of subsequent image features. Then, this first image sample... Divide the material proportionally, specifically using a division formula. The first sub-image sample was obtained. Second sub-image sample and respectively the first sub-image samples Send to the low-rank self-attention branch, and take the second sub-image sample. Send to the standard residual component branch.
[0064] Next, exemplarily, such as Figure 5b The diagram shown is a schematic representation of the internal structure of a low-rank self-attention branch provided in an embodiment of this application. Figure 5b In the internal structure of the low-rank self-attention branch shown, the preset dimensionality reduction matrix may include the query matrix. Key matrix Sum matrix , , , Represents the first sub-image sample The number of channels, Represents the fourth sub-image sample The dimension. Combined Figure 5b The electronic device employs a low-rank self-attention branch, first applying it to the first sub-image sample. Dimensionality reduction can be performed using the second dimensionality reduction formula. The third sub-image sample was obtained. Then, based on the preset dimensionality reduction matrix, the third sub-image sample... The fourth sub-image sample is obtained by mapping, which can be done using the mapping formula. , and Three fourth sub-image samples were obtained, namely the fourth sub-image samples ( ), fourth sub-image sample ( ) and the fourth sub-image sample ( Furthermore, the long-distance dependencies between features in these three fourth sub-image samples are captured, specifically using the attention calculation formula. The fifth sub-image sample was obtained. ,in, This represents a mathematical function that transforms a real-valued vector matrix into a probability distribution; further, it applies to the fifth sub-image sample. To achieve dimensionality increase, the first dimensionality increase formula can be used. The sixth sub-image sample was obtained. .
[0065] Meanwhile, the electronic device employs a standard residual component branch to perform processing on the second sub-image sample. Dimensionality reduction can be achieved using the third dimensionality reduction formula. The seventh sub-image sample was obtained. Then, regarding the seventh sub-image sample... Local feature extraction can be performed, specifically using feature extraction formulas. The eighth sub-image sample was obtained. ,in, This indicates a convolution operation with a 3×3 kernel; and it applies to the eighth sub-image sample. To achieve dimensionality increase, the second dimensionality increase formula can be used. The ninth sub-image sample was obtained. .
[0066] Finally, the electronic device uses the first stitching module to stitch the sixth sub-image sample. and the ninth sub-image sample Then, stitch the components along the channel dimension, specifically using the first stitching formula. This yields the stitched sub-image samples, where... This indicates a stitching operation, and a dimensionality-upgrading module is used to upgrade the dimensionality of the stitched sub-image sample to obtain the first feature image sample. The dimension-up module is a convolutional module with a 1×1 kernel. The above applies to the first stitched image sample. The processing can reduce computational complexity while maintaining the self-attention mechanism's ability to capture dependencies between features, enriching feature representations, and quickly obtaining first feature image samples with high completeness and accuracy. .
[0067] It should be noted that the electronic device determines the first sub-image sample. Second sub-image sample The timing is not limited; the electronic device determines the sixth sub-image sample. and the ninth sub-image sample The timing is not limited.
[0068] Step 203: Using the second image processing module, the resolution of the first feature image sample is improved and then stitched with the third image sample to be tested to obtain the second stitched image sample; and DC-AFM is used to extract, filter and enhance the detailed features of the second stitched image sample to obtain the second feature image sample.
[0069] In step 203, the electronic device uses a second image processing module to improve the resolution of the first feature image sample, making the detailed features in the first feature image sample richer and clearer. Then, the first feature image sample with improved resolution is stitched together with the third test image sample to obtain a second stitched image sample. Since DC-AFM is an improved module that combines multi-scale feature extraction and attention mechanism fusion, aiming to enhance feature extraction and learning capabilities, the electronic device can use DC-AFM to extract, filter, and enhance the detailed features of the second stitched image sample, quickly and accurately obtaining the second feature image sample.
[0070] It should be noted that DC-AFM is an improved module that integrates multi-scale feature extraction and attention mechanisms. Its principle is to effectively extract multi-scale features of images by using dilated convolutions with different dilation rates (such as d=1, d=3, d=5), which can meet the detection needs of targets of different sizes. This allows the crack detection model to take into account both the detailed features and overall structural features of the target in complex scenes, thereby improving the richness of feature representation. In addition, it integrates a channel attention mechanism consisting of global average pooling (AvgPool) and Conv operation and a spatial attention mechanism consisting of Conv operation and nonlinear activation function (such as Sigmoid function) to filter and enhance features from both channel and spatial dimensions. This allows the crack detection model to more accurately locate and identify targets, suppress redundant information, and improve the specificity and effectiveness of features.
[0071] The following section details how the electronic device uses a second image processing module to improve the resolution of the first feature image sample and then stitches it with the third image sample to obtain a second stitched image sample: Optionally, such as Figure 6 The diagram shown is a structural schematic of the second image processing module provided in an embodiment of this application. Figure 6 In this process, the second image processing module 104 may include: a second upsampling module and a second image stitching module; the electronic device uses the second image processing module to improve the resolution of the first feature image sample and then stitches it with the third image sample to be tested to obtain a second stitched image sample, which may include: the electronic device uses the second upsampling module to improve the resolution of the first feature image sample to obtain a second resolution image sample; and uses the second image stitching module to stitch the second resolution image sample and the third image sample to be tested to obtain a second stitched image sample.
[0072] In determining the second stitched image sample, the electronic device can first use a second upsampling module to increase the resolution of the first feature image sample, making the detailed features in the first feature image sample richer and clearer, thus obtaining a second resolution image sample. Then, a second image stitching module is used to stitch the second resolution image sample and the third image sample to be tested to obtain more comprehensive feature information, resulting in a second stitched image sample that is information-rich and highly accurate.
[0073] The following section elaborates on the use of DC-AFM to extract, filter, and enhance the detailed features of the second stitched image sample on an electronic device, resulting in a second feature image sample: In some embodiments, such as Figure 7 The diagram shown is a schematic representation of the DC-AFM provided in an embodiment of this application. Figure 7In this system, the DC-AFM105 may include: a second convolutional module, dilated convolutional modules with different dilation rates, a second stitching module, a global average pooling module, a third convolutional module, a fourth convolutional module, a fusion module, and a fifth convolutional module. The electronic device uses DC-AFM to extract, filter, and enhance detailed features from the second stitched image samples to obtain second feature image samples. This may include: the electronic device using the second convolutional module to reduce the dimensionality of the second feature image samples to obtain second image samples; and the electronic device using dilated convolutional modules with different dilation rates to extract features from the second image samples to obtain multi-scale image samples, where one dilation rate corresponds to one scale. The electronic device employs a second stitching module to stitch together multi-scale image samples to obtain stitched sub-image samples. It also employs a global average pooling module and a third convolution module to generate channel attention weights based on a first activation function and the stitched sub-image samples. Furthermore, it uses a fourth convolution module to generate spatial attention weights based on a second activation function and the stitched sub-image samples. Finally, it employs a fusion module to fuse the stitched sub-image samples, channel attention weights, and spatial attention weights to obtain fused feature image samples. Finally, it uses a fifth convolution module to upscale the fused feature image samples to obtain second feature image samples.
[0074] Among them, the aforementioned second stitched image sample , Represents a three-dimensional tensor; Indicates the second stitched image sample The number of channels, i.e., the length; Indicates the second stitched image sample The number of rows, i.e., the height; Indicates the second stitched image sample The number of columns, i.e., the width.
[0075] In determining the first feature image sample, the electronic device first uses a second convolution module with a 1×1 kernel to process the two stitched image samples. Dimensionality reduction can be achieved using the fourth dimensionality reduction formula. The second image sample was obtained. ,in, This indicates the first convolutional module with a 1×1 kernel, i.e., a convolution operation with a 1×1 kernel; , Indicates the second image sample The number of channels. This can effectively reduce the amount of computation required for subsequent image feature calculations.
[0076] Then, the electronic device uses dilated convolution modules with different dilation rates (e.g., d=1, d=3, d=5) to process the second image sample. Feature extraction is performed to obtain multi-scale image samples. Specifically, a multi-scale feature extraction formula can be used. Three scale image samples were obtained. ,in, This indicates a dilated convolution operation. It's important to note that dilated convolution modules with different dilation rates can expand the receptive field through interval sampling. For example, when d=1, it's similar to ordinary convolution, focusing on local features; when d=3 and d=5, the receptive field increases, capturing a wider range of contextual information. Parallel convolution with different dilation rates allows for the extraction of multi-scale features. Small-scale features focus on detailed features, while large-scale features focus on overall structural features, resulting in image samples at three scales with high accuracy. , .
[0077] Next, the electronic device uses a second stitching module to stitch together the multi-scale image samples, that is, the above three scale image samples. For splicing, the second splicing formula can be used. Obtain spliced sub-image samples , The entire stitching operation integrates image features at various scales, covering information at different levels of the image, enabling the crack detection model to learn richer feature representations and providing a comprehensive feature foundation for subsequent target detection tasks.
[0078] Subsequently, the electronic device employs a channel attention mechanism, which utilizes a global average pooling module and a third convolutional module. Specifically, the pooling formula can be used. and the first activation function Obtain the channel attention weights , ,in, This indicates a global average pooling operation. Indicates the spliced sub-image samples The result of performing a global average pooling operation This is the Sigmoid function.
[0079] Meanwhile, this electronic device employs a spatial attention mechanism, specifically a fourth convolutional module, which can be implemented using a second activation function. To obtain spatial attention weights , ,in, This indicates the kernel size of the fourth convolutional module, which is typically 7×7.
[0080] Finally, the electronic device employs a fusion module to perform stitched sub-image sample processing and channel attention weighting. Spatial attention weights To perform fusion and filter key features, a fusion formula can be used. , to obtain fused feature image samples ,in, This represents element-wise multiplication; then, the fifth convolution module is used to fuse the feature image samples. Dimensional upscaling is performed to output the final features that integrate multi-scale and attention advantages; specifically, the third dimensionality upscaling formula can be used. The second feature image sample is obtained. , .
[0081] Understandably, in the process of determining the second feature image samples using DC-AFM, electronic devices employ a channel attention mechanism. Average pooling is used to globally average each channel, compressing the spatial dimension while preserving channel statistics. Convolution operations generate channel attention weights by learning parameters. This measures the importance of features in each channel. The entire process can be based on channel attention weights. The algorithm enhances critical channels and suppresses secondary channels, focusing on channel features important to the detection task. Based on a spatial attention mechanism, the convolutional module extracts spatial features and captures the spatial distribution information of the feature map; the sigmoid function normalizes the output of the convolutional module to [0,1], generating spatial attention weights. This indicates the importance of each spatial location in the feature map. The entire process can utilize spatial attention weights. Focusing on the target area while suppressing background interference, DC-AFM effectively extracts second image samples through the aforementioned channel attention and spatial attention mechanisms. The multi-scale features are used to obtain multi-scale image samples; at the same time, the enhanced features are selected and the redundancy is suppressed, which helps the crack detection model to more accurately locate the target of interest in crack detection of airport runways. In addition, since DC-AFM has a relatively simple structure, it can also maintain low computational complexity.
[0082] It should be noted that electronic devices determine channel attention weights. Spatial attention weights The timing is not limited.
[0083] Step 204: Using the third image processing module, based on the first image sample to be tested, the first feature image sample, and the second feature image sample, determine the first target feature image sample corresponding to the first image sample to be tested, the second target feature image sample of the second image sample to be tested, and the third target feature image sample of the third image sample to be tested.
[0084] It should be noted that the first target feature image sample, the second target feature image sample, and the third target feature image sample are three predicted crack identification results at three different scales. Specifically, the first target feature image sample is the first predicted crack identification result; the second target feature image sample is the second predicted crack identification result; and the third target feature image sample is the third predicted crack identification result.
[0085] The following section elaborates on how the electronic device uses a third image processing module to determine the first target feature image sample corresponding to the first image sample under test, the second target feature image sample of the second image sample under test, and the third target feature image sample of the third image sample under test, based on the first image sample under test, the first feature image sample, and the second feature image sample: In some embodiments, such as Figure 8 The diagram shown is a structural schematic of the third image processing module provided in an embodiment of this application. Figure 8 In this context, the third image processing module 106 may include: a second LSAF module, a first detection head, a first sub-processing module, a third LSAF module, a second detection head, a second sub-processing module, a fourth LSAF module, and a third detection head. The electronic device employs the third image processing module to determine, based on the first image sample to be tested, the first feature image sample, and the second feature image sample, a first target feature image sample corresponding to the first image sample to be tested, a second target feature image sample of the second image sample, and a third target feature image sample of the third image sample. This may include: the electronic device employing the second LSAF module to capture the long-distance dependencies between features in the second feature image sample to obtain the third feature image sample; the electronic device employing the first detection head to determine the first target feature image sample based on the third feature image sample; the electronic device... The device employs a first sub-processing module to extract features from a first target feature image sample and then stitches it with the first feature image sample to obtain a third stitched image sample. The device also employs a third LSAF module to capture long-distance dependencies between features in the third stitched image sample, resulting in a fourth feature image sample. Finally, the device uses a second detection head to determine a second target feature image sample based on the fourth feature image sample. The device further employs a second sub-processing module to extract features from the fourth feature image sample and then stitches it with a first test image sample to obtain a fourth stitched image sample. The device also employs a fourth LSAF module to capture long-distance dependencies between features in the fourth stitched image sample, resulting in a fifth feature image sample. Finally, the device uses a third detection head to determine a third target feature image sample based on the fifth feature image sample.
[0086] Optionally, the first sub-processing module may include a sixth sub-convolution module and a third stitching module.
[0087] Optionally, the second sub-processing module may include a seventh sub-convolution module and a fourth stitching module.
[0088] In the process of determining the first predicted crack identification result, the electronic device uses a second LSAF module to capture the long-distance dependency between features in the second feature image sample to obtain a third feature image sample. Then, the first detection head is used to map the features in the third feature image sample to the first target feature image sample (such as the first predicted bounding box and / or the first predicted crack classification result).
[0089] In determining the second predicted crack identification result, the electronic device uses a sixth sub-convolution module to extract features from the third feature image sample, and then uses a third stitching module to stitch the extracted third feature image sample with the first feature image sample to obtain a third stitched image sample. Then, a third LSAF module is used to capture the long-distance dependencies between features in the third stitched image sample to obtain a fourth feature image sample. Finally, a second detection head is used to map the features in the fourth feature image sample to a second target feature image sample (such as a second predicted bounding box and / or a second predicted crack classification result).
[0090] In determining the third predicted crack identification result, the electronic device uses a seventh sub-convolution module to extract features from the fourth feature image sample, and then uses a fourth stitching module to stitch the extracted features from the fourth feature image sample with the first test image sample to obtain a fourth stitched image sample. Then, a fourth LSAF module is used to capture the long-distance dependencies between features in the fourth stitched image sample to obtain a fifth feature image sample. Finally, a third detection head is used to map the features in the fifth feature image sample to a third target feature image sample (such as the third predicted bounding box and / or the third predicted crack classification result).
[0091] The first, second, and third predicted bounding boxes are rectangular.
[0092] It should be noted that the second, third, and fourth LSAF modules mentioned above have the same function and structure as the first LSAF module mentioned above, and the calculation process is similar. That is, only the input data and output data are different, which will not be elaborated here.
[0093] Step 205: Update the parameters of the crack detection model based on the first target feature image sample, the second target feature image sample, the third target feature image sample, and the actual crack feature image sample to obtain the trained crack detection model.
[0094] Optionally, the actual crack feature image sample may include: actual bounding boxes and / or actual crack classification results. The actual bounding boxes are rectangular.
[0095] The following section elaborates on how electronic devices update the parameters of the crack detection model to obtain the trained crack detection model: In some embodiments, the electronic device updates the parameters of the crack detection model based on a first target feature image sample, a second target feature image sample, a third target feature image sample, and an actual crack feature image sample to obtain a trained crack detection model. This may include: the electronic device determining a first predicted bounding box in the first target feature image sample, a second predicted bounding box in the second target feature image sample, a third predicted bounding box in the third target feature image sample, and an actual bounding box in the actual crack feature image sample; the electronic device determining a first position-enhanced complete intersection-union ratio (CUIR) based on a first intersection area and a first union area between the first predicted bounding box and the actual bounding box, combined with the first diagonal coordinate information of the first predicted bounding box and the diagonal coordinate information corresponding to the first diagonal coordinate information in the actual bounding box; the electronic device further determines the first position-enhanced complete intersection-union ratio (CUIR) based on the first intersection area and the first union area between the first predicted bounding box and .... Based on the second intersection area and the second union area between the second predicted bounding box and the actual bounding box, combined with the second diagonal coordinate information of the second predicted bounding box and the diagonal coordinate information corresponding to the second diagonal coordinate information in the actual bounding box, the second position-enhanced complete intersection-union ratio is determined. Based on the third intersection area and the third union area between the third predicted bounding box and the actual bounding box, combined with the third diagonal coordinate information of the third predicted bounding box and the diagonal coordinate information corresponding to the third diagonal coordinate information in the actual bounding box, the third position-enhanced complete intersection-union ratio is determined. Based on the first position-enhanced complete intersection-union ratio, the second position-enhanced complete intersection-union ratio, and the third position-enhanced complete intersection-union ratio, the electronic device updates the parameters of the crack detection model, that is, updates the parameters of the multi-scale image determination module 101, and obtains the trained crack detection model.
[0096] During the training of the crack detection model, electronic devices can first determine the first predicted bounding box in the first target feature image sample, the second predicted bounding box in the second target feature image sample, the third predicted bounding box in the third target feature image sample, and the actual bounding box in the actual crack feature image sample.
[0097] Then, for the first target feature image sample, the first intersection area and the first union area between the first predicted bounding box and the actual bounding box are determined, as well as the first diagonal coordinate information of the first predicted bounding box. The first intersection area, the first union area, the first diagonal coordinate information, and the diagonal coordinate information in the actual bounding box corresponding to the first diagonal coordinate information are then calculated to obtain the first position-enhanced complete intersection-union ratio.
[0098] For the second target feature image sample, determine the second intersection area and the second union area between the second predicted bounding box and the actual bounding box, as well as the second diagonal coordinate information of the second predicted bounding box. Then, calculate the second intersection area, the second union area, the second diagonal coordinate information, and the diagonal coordinate information in the actual bounding box corresponding to the second diagonal coordinate information to obtain the second position-enhanced complete intersection-union ratio.
[0099] For the third target feature image sample, determine the third intersection area and the third union area between the third predicted bounding box and the actual bounding box, as well as the third diagonal coordinate information of the third predicted bounding box. Then, calculate the third intersection area, the third union area, the third diagonal coordinate information, and the diagonal coordinate information corresponding to the third diagonal coordinate information in the actual bounding box to obtain the third position-enhanced complete intersection-union ratio.
[0100] At this point, the electronic device obtains three position-enhanced complete cross-union ratios, namely the first position-enhanced complete cross-union ratio, the second position-enhanced complete cross-union ratio, and the third position-enhanced complete cross-union ratio. Then, based on these three position-enhanced complete cross-union ratios, the parameters of the crack detection model are updated to obtain the trained crack detection model.
[0101] The following section elaborates on how electronic devices determine the first position-enhanced perfect intersection-union ratio based on the first intersection area and the first union area between the first predicted bounding box and the actual bounding box, combined with the first diagonal coordinate information of the first predicted bounding box and the diagonal coordinate information corresponding to the first diagonal coordinate information in the actual bounding box: In some embodiments, in the first case: when the first diagonal coordinate information is the first upper left corner coordinate information and the first lower right corner coordinate information of the first predicted bounding box, the diagonal coordinate information corresponding to the first diagonal coordinate information is the second upper left corner coordinate information and the second lower right corner coordinate information of the actual bounding box.
[0102] Based on this, the electronic device determines the first position-enhanced complete intersection-union ratio (CUIR) by combining the first intersection area and the first union area between the first predicted bounding box and the actual bounding box, along with the first diagonal coordinate information of the first predicted bounding box and the diagonal coordinate information of the actual bounding box corresponding to the first diagonal coordinate information. This can include: in a first case, the electronic device determines the first distance information based on the first upper-left corner coordinate information, the second upper-left corner coordinate information, and the length and width information of the target bounding box, where the target bounding box is either the first predicted bounding box or the actual bounding box; determines the second distance information based on the first lower-right corner coordinate information, the second lower-right corner coordinate information, and the length and width information; and determines the first position-enhanced complete intersection-union ratio based on the interaction ratio of the first intersection area and the first union area, the first distance information, and the second distance information.
[0103] In determining the enhanced perfect cross-union ratio at the first location, the electronic device assumes that the area of the first prediction box is... The actual area of the predicted bounding box is At this point, the interaction ratio formula can be used to determine the interaction ratio between the area of the first intersection and the area of the first union. .
[0104] Specifically, the above interaction ratio formula is as follows: ,in, Indicates the area of the first intersection. This represents the area of the first union.
[0105] Then, to more accurately measure the overlap of the bounding boxes, let's assume the coordinates of the first top-left corner of the first predicted bounding box are... The first bottom right corner coordinate information is The actual coordinates of the second top-left corner of the bounding box are: Second, bottom right corner coordinate information The length information of the first predicted bounding box or the actual bounding box is: The altitude information is At this point, the first distance formula can be used to calculate the first distance information. The second distance information is calculated using the second distance formula. .
[0106] Specifically, the formula for the first distance is: ; The second distance formula is: .
[0107] Finally, the electronic device uses the PECIoU loss function to calculate the first position-enhanced complete crossover ratio.
[0108] Specifically, the expression for the PECIoU loss function is: PECIoU1 ,in, This represents the first parameter that allows the crack detection model to be trained. This represents the second parameter that allows the crack detection model to be trained. As can be seen from the expression of the PECIoU loss function, PECIoU1 can better locate the position of the crack prediction box.
[0109] In some embodiments, the second case is: when the first diagonal coordinate information is the first upper right corner coordinate information and the first lower left corner coordinate information of the first predicted bounding box, the diagonal coordinate information corresponding to the first diagonal coordinate information is the second upper right corner coordinate information and the second lower left corner coordinate information of the actual bounding box.
[0110] Based on this, the electronic device determines the first position-enhanced complete intersection-union ratio (CUIR) by combining the first intersection area and the first union area between the first predicted bounding box and the actual bounding box, along with the first diagonal coordinate information of the first predicted bounding box and the diagonal coordinate information of the actual bounding box corresponding to the first diagonal coordinate information. This can include: in a second case, the electronic device determines a third distance information based on the first upper right corner coordinate information, the second upper right corner coordinate information, and length and width information; determines a fourth distance information based on the first lower left corner coordinate information, the second lower left corner coordinate information, and length and width information; and determines the first position-enhanced complete intersection-union ratio based on the intersection ratio, the third distance information, and the fourth distance information.
[0111] In determining the first position enhanced perfect intersection-union ratio, the electronic device can use the above-mentioned interaction ratio formula to determine the interaction ratio between the first intersection area and the first union area. .
[0112] Then, to more accurately measure the overlap of the bounding boxes, let's assume the coordinates of the first upper right corner of the first predicted bounding box are... The first bottom left corner coordinate information is The actual bounding box's second upper right corner coordinates are: Second, bottom left corner coordinate information At this point, the third distance formula can be used to calculate the first distance information. The second distance information is calculated using the fourth distance formula. .
[0113] Specifically, the formula for the first distance is: ; The second distance formula is: .
[0114] Finally, the electronic device uses the aforementioned PECIoU loss function to calculate the first position-enhanced complete crossover ratio.
[0115] Understandably, in both the first and second scenarios, the electronic device determines PECioU1 by using the PECioU loss function as the bounding box regression loss to assist the crack detection model in calculating the bounding box, aiming to improve the performance of object detection and image segmentation tasks. It should be noted that this PECioU loss function proposes a bounding box similarity comparison index based on minimum point distance. This index comprehensively considers factors such as overlapping or non-overlapping regions, center point distance, width and height deviations, etc., to comprehensively measure the similarity between the first predicted bounding box and the actual bounding box, making the crack detection model converge faster and the regression results more accurate, thus effectively improving the performance of object detection tasks.
[0116] Furthermore, the calculation process for the second-position enhanced complete crossover ratio PECIoU2 and the third-position enhanced complete crossover ratio PECIoU3 is similar to that of PECIoU1 mentioned above, and will not be elaborated here.
[0117] In some embodiments, the electronic device updates the parameters of the crack detection model based on the first position-enhanced complete intersection-union ratio (CUIR), the second position-enhanced CUIR, and the third position-enhanced CUIR to obtain a trained crack detection model. This may include: the electronic device updating the parameters of the crack detection model based on the first position-enhanced CUIR, the second position-enhanced CUIR, and the third position-enhanced CUIR to obtain a current crack detection model; the electronic device records the number of training iterations of the current crack detection model; when the number of training iterations reaches the maximum number of iterations, and the difference between the first position-enhanced CUIR, the second position-enhanced CUIR, and the third position-enhanced CUIR and the corresponding previous position-enhanced CUIR is less than a preset difference threshold, the electronic device determines the current crack detection model as the trained crack detection model; when the number of training iterations does not reach the maximum number of iterations, and / or the difference is greater than or equal to the preset difference threshold, the electronic device continues to train the current crack detection model until the trained crack detection model is determined.
[0118] The electronic device is pre-configured with the network environment of the crack detection module (e.g., Python version 3.12, deep learning framework PyTorch 2.6.0, and CUDA 12.6 for training acceleration), and the initial learning rate is set to 0.01. In addition, the number of images input to the crack detection model in each batch is set to 8. Then, the crack detection model is trained without using pre-trained weights. After each cycle of the training process, the overall network loss is calculated and the model is iterated. Specifically, the parameters of the crack detection model are updated based on the enhanced complete intersection-union ratios (ECUs) at the three positions mentioned above to obtain the current crack detection model. The number of training iterations of the current crack detection model is recorded. If the number of training iterations reaches the maximum number of iterations (e.g., 50), and the difference between the enhanced ECUs at the three positions and the corresponding previous enhanced ECU is less than a preset difference threshold, then the training of the current crack detection model is stopped after determining that its performance is stable, and the current crack detection model is determined as the trained crack detection model. If the number of training iterations is less than or equal to the maximum number of iterations, and the difference between the enhanced ECUs at the three positions and the corresponding previous enhanced ECU is less than a preset difference threshold, then the performance of the current crack detection model is determined to be unstable, and the current crack detection model is trained again until the trained crack detection model is determined.
[0119] In this embodiment, the technical solutions of steps 201-205 employ an LSAF module in the crack detection model. This module uses a low-rank self-attention mechanism to capture long-distance dependencies and standard convolution to extract local features. The combination of these two approaches allows the LSAF module to simultaneously learn the long-distance relationships and local details of features, enriching feature representation. Furthermore, the low-rank attention mechanism reduces the computational load and memory usage of self-attention through dimensionality reduction, thereby lowering computational complexity and effectively improving the detection efficiency of crack identification results for airport runways. In addition, adding DC-AFM to the crack detection model allows for the effective acquisition of multi-scale image features by using dilated convolutions with different dilation rates. This meets the detection requirements for targets of different sizes, enabling the crack detection model to consider both the detailed features and overall structural features of targets in complex scenes, thus enhancing the richness of feature representation. Simultaneously, by filtering and enhancing features from both channel and spatial dimensions, redundant information is suppressed, allowing the crack detection model to more accurately locate and identify targets, further improving feature targeting and effectiveness, and ultimately enhancing the detection accuracy of crack identification results. In other words, the entire method can train a crack detection model with high accuracy, which can effectively improve the accuracy and speed of crack identification on airport runways. This crack detection model can be used in the future to quickly and accurately detect cracks on airport runways.
[0120] To better understand the embodiments of this application, the crack detection model is illustrated below with examples: For example, such as Figure 9 The diagram shown is a structural connection schematic of the crack detection model provided in an embodiment of this application. Figure 9 The crack detection model may include: a multi-scale image determination module 101, a first image processing module 102, a first LSAF module 103, a second image processing module 104, a DC-AFM module 105, and a third image processing module 106.
[0121] The multi-scale image determination module 101 may include: a first sub-Conv module, a second sub-Conv module, a first C3k2 module, a third sub-Conv module, a second C3k2 module, a fourth sub-Conv module, a third C3k2 module, a fifth sub-Conv module, a fourth C3k2 module, a C2PSA module, and an SPFF module.
[0122] The first image processing module 102 may include a first upsample module and a first image concat module.
[0123] The second image processing module 104 may include: a second Upsample module and a second image Concat module.
[0124] The third image processing module 106 may include: a second LSAF module, a first head, a sixth sub-Conv module, a third Concat module, a third LSAF module, a second head, a seventh sub-Conv module, a fourth Concat module, a fourth LSAF module, and a third head.
[0125] It should be noted that, Figure 9 The functions of each module in the crack detection model shown are as follows: Figures 3-8 The corresponding modules shown have the same function, so they will not be described in detail here.
[0126] The following example illustrates the training method for crack detection models: For example, such as Figure 10 The diagram shown is a flowchart illustrating the crack detection model training method provided in an embodiment of this application. Figure 10 The method described is as follows: First, a labeled airport runway crack dataset is constructed, and then this dataset is divided into a training set and a test set. Next, the crack detection model is initialized and trained using the training set to update its parameters. The current training iteration count and the augmented complete intersection-union ratio (PCU) at the three locations are recorded. Then, if the training iteration count reaches the maximum, and the differences between the PCU at these three locations and the previous PCU are all less than a preset threshold, the model training is stopped once the current crack detection model is deemed stable. If the training iteration count is less than or equal to the maximum, and the differences between the PCU at these three locations and the previous PCU are less than a preset threshold, the current crack detection model is deemed unstable, and training continues until a suitable crack detection model is determined. Finally, the test set is input into the trained crack detection model, and the crack identification results are output.
[0127] It should be noted that the entity executing the crack detection method provided in this application embodiment can be a crack detection device or an electronic device, and no specific limitation is made here.
[0128] Optionally, the aforementioned electronic devices may include computers, mobile terminals, and wearable devices, etc.
[0129] The crack detection method provided in this application embodiment will be described in detail below using an electronic device as an example: like Figure 11 The diagram shown is a schematic flowchart of the crack detection method provided in an embodiment of this application. Figure 11 The method includes the following steps 1101-1102.
[0130] Step 1101: Obtain the image of the airport runway to be tested.
[0131] The image to be tested is data from the test set.
[0132] Step 1102: Input the image to be tested into the crack detection model to obtain the crack recognition result output by the crack detection model. The crack recognition result includes the crack bounding box and crack classification result.
[0133] Among them, the crack detection model is based on, for example Figure 2 The crack detection model is obtained after training using the aforementioned training method.
[0134] For example, such as Figure 12a The image shown is a schematic diagram of a target feature image provided in an embodiment of this application. Figure 12a It can be seen that there are four cracks in the airport runway. For example... Figure 12b The image shown is a schematic diagram of a target feature image provided in an embodiment of this application. Figure 12b It can be seen that there are six cracks in the airport runway.
[0135] In this embodiment of the application, the technical solution of steps 1101-1102 adopts a crack detection model, which can quickly and accurately detect the crack identification results of the airport runway.
[0136] The crack detection model training device provided in the embodiments of this application is described below. The crack detection model training device described below can be referred to in correspondence with the crack detection model training method described above.
[0137] Figure 13 This is a schematic diagram of the crack detection model training device provided in an embodiment of this application. Figure 13 As shown, the crack detection model training device may include: a multi-scale image determination module 101, a first image processing module 102, a first low-rank self-attention fusion LSAF module 103, a second image processing module 104, a dilated convolution-attention fusion module DC-AFM 105, a third image processing module 106, and a parameter update module 107.
[0138] The multi-scale image determination module 101 is used to determine a first image sample, a second image sample, and a third image sample based on the initial image sample to be tested of the airport runway sample. The first image processing module 102 is used to improve the resolution of the first image sample to be tested and then stitch it with the second image sample to be tested to obtain the first stitched image sample. LSAF module 103 is used to capture the long-distance dependencies between features in the first stitched image sample to obtain the first feature image sample; The second image processing module 104 is used to improve the resolution of the first feature image sample and then stitch it with the third image sample to be tested to obtain a second stitched image sample. The DC-AFM105 is used to extract, filter, and enhance the detailed features of the second stitched image sample to obtain the second feature image sample. The third image processing module 106 is used to determine, based on the first image sample to be tested, the second target feature image sample of the second image sample to be tested, and the third target feature image sample of the third image sample to be tested, the first feature image sample, and the second feature image sample, the first target feature image sample and the second feature image sample, the first target feature image sample and the second feature image sample, the first target feature image sample and the second feature image sample, the first target feature image sample and the second feature image sample, the first target feature image sample and the second target ... The parameter update module 107 is used to update the parameters of the crack detection model based on the first target feature image sample, the second target feature image sample, the third target feature image sample and the actual crack feature image sample, so as to obtain the trained crack detection model.
[0139] Optionally, the LSAF module 103 includes: a first convolution module, a low-rank self-attention branch, a standard residual component branch, a first stitching module, and a dimensionality-upgrading module; wherein, the first convolution module is used to reduce the dimensionality of the first stitched image sample to obtain a first image sample, and to split the first image sample to obtain a first sub-image sample and a second sub-image sample. The low-rank self-attention branch is used to reduce the dimensionality of the first sub-image sample to obtain the third sub-image sample; the third sub-image sample is mapped according to the preset dimensionality reduction matrix to obtain the fourth sub-image sample; the long-distance dependencies between features in the fourth sub-image sample are captured to obtain the fifth sub-image sample; and the fifth sub-image sample is increased in dimensionality to obtain the sixth sub-image sample. The standard residual component branch is used to reduce the dimensionality of the second sub-image sample to obtain the seventh sub-image sample; to extract local features from the seventh sub-image sample to obtain the eighth sub-image sample; and to increase the dimensionality of the eighth sub-image sample to obtain the ninth sub-image sample. The first stitching module is used to stitch the sixth sub-image sample and the ninth sub-image sample together; The dimension-up module is used to increase the dimension of the stitched sub-image samples to obtain the first feature image sample.
[0140] Optionally, the DC-AFM105 includes: a second convolution module, dilated convolution modules with different dilation rates, a second stitching module, a global average pooling module, a third convolution module, a fourth convolution module, a fusion module, and a fifth convolution module; wherein, the second convolution module is used to reduce the dimensionality of the second feature image sample to obtain the second image sample; Dilated convolutional modules with different dilation rates are used to extract features from the second image sample to obtain multi-scale image samples, with one dilation rate corresponding to one scale image sample. The second stitching module is used to stitch together multi-scale image samples to obtain stitched sub-image samples; The global average pooling module and the third convolution module are used to generate channel attention weights based on the first activation function and the stitched sub-image samples; The fourth convolutional module is used to generate spatial attention weights based on the second activation function and the stitched sub-image samples; The fusion module is used to fuse the stitched sub-image samples, channel attention weights, and spatial attention weights to obtain fused feature image samples; The fifth convolutional module is used to increase the dimensionality of the fused feature image samples to obtain the second feature image samples.
[0141] Optionally, the parameter update module 107 is specifically used to determine the first predicted bounding box in the first target feature image sample, the second predicted bounding box in the second target feature image sample, the third predicted bounding box in the third target feature image sample, and the actual bounding box in the actual crack feature image sample; based on the first intersection area and the first union area between the first predicted bounding box and the actual bounding box, combined with the first diagonal coordinate information of the first predicted bounding box and the diagonal coordinate information corresponding to the first diagonal coordinate information in the actual bounding box, determine the first position-enhanced complete intersection-union ratio; based on the second intersection area and the second union area between the second predicted bounding box and the actual bounding box, determine the first position-enhanced complete intersection-union ratio. The area of the first intersection and the area of the third union between the predicted and actual bounding boxes are used to determine the second position-enhanced complete intersection-union ratio (CUIR). Based on the area of the third intersection and the area of the third union between the predicted and actual bounding boxes, and combined with the area of the third diagonal coordinates of the predicted and actual bounding boxes, the area of the third position-enhanced complete intersection-union ratio (CUIR) is determined. Based on the CUIRs of the first, second, and third positions, the parameters of the crack detection model are updated to obtain the trained crack detection model.
[0142] Optionally, in the first case: when the first diagonal coordinate information is the first upper-left corner coordinate information and the first lower-right corner coordinate information of the first predicted bounding box, the diagonal coordinate information corresponding to the first diagonal coordinate information is the second upper-left corner coordinate information and the second lower-right corner coordinate information of the actual bounding box; in the second case: when the first diagonal coordinate information is the first upper-right corner coordinate information and the first lower-left corner coordinate information of the first predicted bounding box, the diagonal coordinate information corresponding to the first diagonal coordinate information is the second upper-right corner coordinate information and the second lower-left corner coordinate information of the actual bounding box. The parameter update module 107 is specifically used to: in the first case, determine the first distance information based on the first upper-left corner coordinate information, the second upper-left corner coordinate information, and the length and width information of the target bounding box, wherein the target bounding box is either the first predicted bounding box or the actual bounding box; determine the second distance information based on the first lower-right corner coordinate information, the second lower-right corner coordinate information, and the length and width information; and determine the first position-enhanced complete intersection-union ratio based on the interaction ratio of the first intersection area and the first union area, the first distance information, and the second distance information. In the second case, determine the third distance information based on the first upper-right corner coordinate information, the second upper-right corner coordinate information, and the length and width information; determine the fourth distance information based on the first lower-left corner coordinate information, the second lower-left corner coordinate information, and the length and width information; and determine the first position-enhanced complete intersection-union ratio based on the interaction ratio, the third distance information, and the fourth distance information.
[0143] Optionally, the parameter update module 107 is specifically used to update the parameters of the crack detection model based on the first position enhanced complete intersection-union ratio, the second position enhanced complete intersection-union ratio, and the third position enhanced complete intersection-union ratio to obtain the current crack detection model; record the training number of the current crack detection model; when the training number reaches the maximum number of iterations, and the difference between the first position enhanced complete intersection-union ratio, the second position enhanced complete intersection-union ratio, and the third position enhanced complete intersection-union ratio and the corresponding previous position enhanced complete intersection-union ratio is less than a preset difference threshold, the current crack detection model is determined as the trained crack detection model; when the training number does not reach the maximum number of iterations, and / or the difference is greater than or equal to the preset difference threshold, the current crack detection model continues to be trained until the trained crack detection model is determined.
[0144] Optionally, the third image processing module 106 includes: a second LSAF module, a first detection head, a first sub-processing module, a third LSAF module, a second detection head, a second sub-processing module, a fourth LSAF module, and a third detection head; wherein, the second LSAF module is used to capture the long-distance dependencies between features in the second feature image sample to obtain the third feature image sample. The first detection head is used to determine the first target feature image sample based on the third feature image sample; The first sub-processing module is used to extract features from the first target feature image sample and then stitch it with the first feature image sample to obtain the third stitched image sample. The third LSAF module is used to capture the long-distance dependencies between features in the third stitched image sample to obtain the fourth feature image sample. The second detection head is used to determine the second target feature image sample based on the fourth feature image sample; The second sub-processing module is used to extract features from the fourth feature image sample and then stitch it together with the first image sample to be tested to obtain the fourth stitched image sample. The fourth LSAF module is used to capture the long-distance dependencies between features in the fourth stitched image sample to obtain the fifth feature image sample; The third detection head is used to determine the third target feature image sample based on the fifth feature image sample.
[0145] The crack detection device provided in the embodiments of this application is described below. The crack detection device described below can be referred to in correspondence with the crack detection method described above.
[0146] Figure 14 This is a schematic diagram of the crack detection device provided in an embodiment of this application. Figure 14 As shown, the crack detection device may include an image acquisition module 1401 and a crack recognition module 1402.
[0147] Image acquisition module 1401 is used to acquire the image of the airport runway to be tested.
[0148] The crack recognition module 1402 is used to input the image to be tested into the crack detection model and obtain the crack recognition result output by the crack detection model. The crack recognition result includes crack bounding boxes and crack classification results; wherein, the crack detection model is based on Figure 2 The crack detection model shown is obtained after training using the training method.
[0149] On the other hand, embodiments of this application also provide a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the crack detection model training method or the crack detection method provided by the above methods.
[0150] In another aspect, embodiments of this application also provide a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the crack detection model training method or the crack detection method provided by the above methods.
[0151] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0152] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0153] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for training a crack detection model, characterized in that, The crack detection model includes: a multi-scale image determination module, a first image processing module, a first low-rank self-attention fusion (LSAF) module, a second image processing module, a dilated convolution-attention fusion (DC-AFM) module, and a third image processing module; the method includes: Using the multi-scale image determination module, a first image sample, a second image sample, and a third image sample are determined based on the initial image samples to be tested of the airport runway sample. The first image processing module is used to improve the resolution of the first image sample to be tested and then stitch it with the second image sample to be tested to obtain a first stitched image sample; and the first low-rank self-attention fusion LSAF module is used to capture the long-distance dependencies between features in the first stitched image sample to obtain a first feature image sample. The second image processing module is used to improve the resolution of the first feature image sample and then stitch it with the third image sample to be tested to obtain a second stitched image sample; and the DC-AFM is used to extract, filter and enhance the detailed features of the second stitched image sample to obtain a second feature image sample. Using the third image processing module, based on the first image sample to be tested, the first feature image sample, and the second feature image sample, a first target feature image sample corresponding to the first image sample to be tested, a second target feature image sample of the second image sample to be tested, and a third target feature image sample of the third image sample to be tested are determined. Based on the first target feature image sample, the second target feature image sample, the third target feature image sample, and the actual crack feature image sample, the parameters of the crack detection model are updated to obtain the trained crack detection model. The step of updating the parameters of the crack detection model based on the first target feature image sample, the second target feature image sample, the third target feature image sample, and the actual crack feature image sample to obtain the trained crack detection model includes: Determine the first predicted bounding box in the first target feature image sample, the second predicted bounding box in the second target feature image sample, the third predicted bounding box in the third target feature image sample, and the actual bounding box in the actual crack feature image sample; Based on the first intersection area and the first union area between the first predicted bounding box and the actual bounding box, and combined with the first diagonal coordinate information of the first predicted bounding box and the diagonal coordinate information of the actual bounding box corresponding to the first diagonal coordinate information, the first position-enhanced perfect intersection-union ratio is determined. Based on the second intersection area and the second union area between the second predicted bounding box and the actual bounding box, and combined with the second diagonal coordinate information of the second predicted bounding box and the diagonal coordinate information of the actual bounding box corresponding to the second diagonal coordinate information, the second position-enhanced perfect intersection-union ratio is determined. Based on the third intersection area and the third union area between the third predicted bounding box and the actual bounding box, and combined with the third diagonal coordinate information of the third predicted bounding box and the diagonal coordinate information in the actual bounding box corresponding to the third diagonal coordinate information, the third position-enhanced complete intersection-union ratio is determined. Based on the enhanced complete intersection-union ratio at the first position, the enhanced complete intersection-union ratio at the second position, and the enhanced complete intersection-union ratio at the third position, the parameters of the crack detection model are updated to obtain the current crack detection model. Record the number of training iterations for the current crack detection model; When the number of training iterations reaches the maximum number of iterations, and the differences between the first position augmented complete intersection-union ratio, the second position augmented complete intersection-union ratio, and the third position augmented complete intersection-union ratio and the corresponding previous position augmented complete intersection-union ratio are less than a preset difference threshold, the current crack detection model is determined as the trained crack detection model. If the number of training iterations does not reach the maximum number of iterations, and / or the difference is greater than or equal to the preset difference threshold, the current crack detection model continues to be trained until the trained crack detection model is determined.
2. The crack detection model training method according to claim 1, characterized in that, The first low-rank self-attention fusion LSAF module includes: a first convolutional module, a low-rank self-attention branch, a standard residual component branch, a first stitching module, and a dimensionality-upgrading module; the first low-rank self-attention fusion LSAF module is used to capture the long-distance dependencies between features in the first stitched image samples to obtain the first feature image samples, including: Using the first convolution module, the first stitched image sample is dimensionality reduced to obtain a first image sample, and the first image sample is split to obtain a first sub-image sample and a second sub-image sample. The first sub-image sample is dimensionality reduced using the low-rank self-attention branch to obtain the third sub-image sample; the third sub-image sample is mapped according to the preset dimensionality reduction matrix to obtain the fourth sub-image sample; the long-distance dependencies between features in the fourth sub-image sample are captured to obtain the fifth sub-image sample; and the fifth sub-image sample is dimensionality increased to obtain the sixth sub-image sample. Using the standard residual component branch, the second sub-image sample is dimensionality reduced to obtain the seventh sub-image sample; local feature extraction is performed on the seventh sub-image sample to obtain the eighth sub-image sample; and the eighth sub-image sample is dimensionality increased to obtain the ninth sub-image sample. The first stitching module is used to stitch the sixth sub-image sample and the ninth sub-image sample together, and the dimensionality-upgrading module is used to upgrade the dimensionality of the stitched sub-image sample to obtain the first feature image sample.
3. The crack detection model training method according to claim 1, characterized in that, The DC-AFM includes: a second convolutional module, dilated convolutional modules with different dilation rates, a second stitching module, a global average pooling module, a third convolutional module, a fourth convolutional module, a fusion module, and a fifth convolutional module; the DC-AFM is used to extract, filter, and enhance detailed features of the second stitched image sample to obtain a second feature image sample, including: The second convolution module is used to reduce the dimensionality of the second feature image sample to obtain the second image sample; Using the dilated convolution modules with different dilation rates, feature extraction is performed on the second image sample to obtain multi-scale image samples, with one dilation rate corresponding to one scale image sample; The second stitching module is used to stitch the multi-scale image samples together to obtain stitched sub-image samples; Using the global average pooling module and the third convolution module, channel attention weights are generated based on the first activation function and the stitched sub-image samples; Using the fourth convolution module, spatial attention weights are generated based on the second activation function and the stitched sub-image samples; The fusion module is used to fuse the stitched sub-image samples, the channel attention weights, and the spatial attention weights to obtain fused feature image samples. The fifth convolution module is used to increase the dimensionality of the fused feature image sample to obtain the second feature image sample.
4. The crack detection model training method according to claim 1, characterized in that, In the first case: when the first diagonal coordinate information is the first upper left corner coordinate information and the first lower right corner coordinate information of the first predicted bounding box, the diagonal coordinate information corresponding to the first diagonal coordinate information is the second upper left corner coordinate information and the second lower right corner coordinate information of the actual bounding box. The second scenario: When the first diagonal coordinate information is the first upper right corner coordinate information and the first lower left corner coordinate information of the first predicted bounding box, the diagonal coordinate information corresponding to the first diagonal coordinate information is the second upper right corner coordinate information and the second lower left corner coordinate information of the actual bounding box. The step of determining the first position-enhanced perfect intersection-union ratio (MIR) based on the first intersection area and the first union area between the first predicted bounding box and the actual bounding box, combined with the first diagonal coordinate information of the first predicted bounding box and the diagonal coordinate information of the actual bounding box corresponding to the first diagonal coordinate information, includes: In the first scenario, a first distance is determined based on the first top-left corner coordinates, the second top-left corner coordinates, and the length and width of the target bounding box, wherein the target bounding box is either the first predicted bounding box or the actual bounding box; a second distance is determined based on the first bottom-right corner coordinates, the second bottom-right corner coordinates, and the length and width; and a first position-enhanced perfect intersection-union ratio is determined based on the interaction ratio of the first intersection area to the first union area, the first distance, and the second distance. In the second case, a third distance information is determined based on the first upper right corner coordinate information, the second upper right corner coordinate information, the length information, and the width information; a fourth distance information is determined based on the first lower left corner coordinate information, the second lower left corner coordinate information, the length information, and the width information; and the first position-enhanced complete intersection-union ratio is determined based on the interaction ratio, the third distance information, and the fourth distance information.
5. The crack detection model training method according to any one of claims 1-3, characterized in that, The third image processing module includes: a second low-rank self-attention fusion LSAF module, a first detection head, a first sub-processing module, a third low-rank self-attention fusion LSAF module, a second detection head, a second sub-processing module, a fourth low-rank self-attention fusion LSAF module, and a third detection head; the step of using the third image processing module to determine, based on the first image sample to be tested, the first feature image sample, and the second feature image sample, the first target feature image sample corresponding to the first image sample to be tested, the second target feature image sample of the second image sample to be tested, and the third target feature image sample of the third image sample to be tested, includes: The second low-rank self-attention fusion LSAF module is used to capture the long-distance dependencies between features in the second feature image sample to obtain the third feature image sample; Using the first detection head, the first target feature image sample is determined based on the third feature image sample; The first sub-processing module is used to extract features from the first target feature image sample and then stitch it with the first feature image sample to obtain a third stitched image sample. The third low-rank self-attention fusion LSAF module is used to capture the long-distance dependencies between features in the third stitched image sample to obtain the fourth feature image sample. Using the second detection head, the second target feature image sample is determined based on the fourth feature image sample; The second sub-processing module is used to extract features from the fourth feature image sample and then stitch it with the first image sample to be tested to obtain the fourth stitched image sample. The fourth low-rank self-attention fusion LSAF module is used to capture the long-distance dependencies between features in the fourth stitched image sample to obtain the fifth feature image sample. Using the third detection head, the third target feature image sample is determined based on the fifth feature image sample.
6. A crack detection method, characterized in that, include: Acquire the image of the airport runway to be tested; The image to be tested is input into the crack detection model to obtain the crack identification result output by the crack detection model. The crack identification result includes: crack bounding box and crack classification result. The crack detection model is obtained by training based on the crack detection model training method described in any one of claims 1-5.
7. A crack detection model training device, characterized in that, include: The system comprises a multi-scale image determination module, a first image processing module, a first low-rank self-attention fusion (LSAF) module, a second image processing module, a dilated convolution-attention fusion (DC-AFM) module, a third image processing module, and a parameter update module. The multi-scale image determination module is used to determine the first image sample, the second image sample, and the third image sample based on the initial image sample to be tested of the airport runway sample. The first image processing module is used to improve the resolution of the first image sample to be tested and then stitch it with the second image sample to be tested to obtain a first stitched image sample. The first low-rank self-attention fusion (LSAF) module is used to capture the long-distance dependencies between features in the first stitched image sample to obtain the first feature image sample. The second image processing module is used to improve the resolution of the first feature image sample and then stitch it with the third image sample to be tested to obtain a second stitched image sample. The DC-AFM is used to extract, filter, and enhance the detailed features of the second stitched image sample to obtain the second feature image sample. The third image processing module is used to determine, based on the first image sample to be tested, the first feature image sample, and the second feature image sample, the first target feature image sample corresponding to the first image sample to be tested, the second target feature image sample corresponding to the second image sample to be tested, and the third target feature image sample corresponding to the third image sample to be tested; The parameter update module is used to update the parameters of the crack detection model based on the first target feature image sample, the second target feature image sample, the third target feature image sample, and the actual crack feature image sample to obtain a trained crack detection model. The step of updating the parameters of the crack detection model based on the first target feature image sample, the second target feature image sample, the third target feature image sample, and the actual crack feature image sample to obtain the trained crack detection model includes: determining a first predicted bounding box in the first target feature image sample, a second predicted bounding box in the second target feature image sample, a third predicted bounding box in the third target feature image sample, and an actual bounding box in the actual crack feature image sample; determining a first position-enhanced complete intersection-union ratio (CUI) based on the first intersection area and the first union area between the first predicted bounding box and the actual bounding box, combined with the first diagonal coordinate information of the first predicted bounding box and the diagonal coordinate information corresponding to the first diagonal coordinate information in the actual bounding box; and determining a second position-enhanced complete intersection-union ratio (CUI) based on the second intersection area and the second union area between the second predicted bounding box and the actual bounding box, combined with the second diagonal coordinate information of the second predicted bounding box. Based on the diagonal coordinates of the actual bounding box corresponding to the second diagonal coordinates, the second enhanced complete intersection-union ratio is determined; based on the third intersection area and third union area between the third predicted bounding box and the actual bounding box, combined with the third diagonal coordinates of the third predicted bounding box and the diagonal coordinates of the actual bounding box corresponding to the third diagonal coordinates, the third enhanced complete intersection-union ratio is determined; based on the first enhanced complete intersection-union ratio, the second enhanced complete intersection-union ratio, and the third enhanced complete intersection-union ratio, the crack detection model is updated with parameters to obtain the current crack detection model. The training process involves: recording the number of training iterations of the current crack detection model; determining the current crack detection model as the trained crack detection model when the number of training iterations reaches the maximum number of iterations and the differences between the first position augmented complete intersection-union ratio, the second position augmented complete intersection-union ratio, and the third position augmented complete intersection-union ratio and the corresponding previous position augmented complete intersection-union ratio are less than a preset difference threshold; and continuing to train the current crack detection model until the trained crack detection model is determined when the number of training iterations does not reach the maximum number of iterations and / or the difference is greater than or equal to the preset difference threshold.
8. A crack detection device, characterized in that, include: Image acquisition module and crack recognition module; The image acquisition module is used to acquire the image of the airport runway to be tested; The crack recognition module is used to input the image to be tested into the crack detection model to obtain the crack recognition result output by the crack detection model. The crack recognition result includes crack bounding boxes and crack classification results. The crack detection model is obtained by training based on the crack detection model training method according to any one of claims 1-5.
Citation Information
Patent Citations
Crack detection method and system based on deep learning
CN116416244A