A lightweight real-time intelligent detection method for road cracks
By designing a lightweight generator network and a self-attention discriminator network, the problem of difficult to balance detection performance, model scale and processing speed in the prior art is solved, and high-performance and high-real-time road crack detection suitable for intelligent patrol vehicles is achieved.
Patent Information
- Application Number
- CN202510369618.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-27
AI Technical Summary
The existing deep learning-based intelligent detection algorithm for road cracks cannot balance detection performance, model scale, and model processing speed, making it difficult to deploy in intelligent patrol vehicles and robots.
A lightweight generator network is designed, including a U-shaped master segmentation network and an adaptive boundary extraction module. Combined with the self-attention discriminator network, the output results are supervised from both global and local levels to achieve high-performance and high-real-time road crack detection.
It achieves a balance of detection performance, model scale and model processing speed, and is suitable for on-board deployment of intelligent patrol vehicles, realizing real-time and accurate fully automated intelligent road crack detection.
Smart Images

Figure CN119888383B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of image processing and artificial intelligence, and in particular to a lightweight real-time intelligent detection method for road cracks. Background Art
[0002] In the context of the development of digital twin technology and smart cities, the intelligent detection of road damage has become an important link in ensuring road quality, safety, traffic efficiency, and driving comfort. If road damage fails to be promptly identified and repaired, it will not only affect the driving experience but also pose a serious threat to vehicle performance, traffic safety, and even life and property safety, and may even cause irreversible personal injuries and economic losses in some cases. At present, the detection of road damage still mainly relies on manual inspections, that is, structural engineers and certified inspectors regularly conduct on-site surveys and record the specific locations of the damage. However, this traditional manual inspection method is easily affected by factors such as driving speed, the experience level of inspectors, and attention concentration, resulting in low detection efficiency, high operating costs, and difficulty in effectively ensuring the safety of inspectors. In addition, due to the high dependence of manual detection on the subjective judgment and experience accumulation of inspectors, the results often lack objectivity and consistency. Therefore, with the rapid development of artificial intelligence technology, some technology companies are actively researching and developing road damage detection technologies based on intelligent algorithms to promote the application of artificial intelligence in the field of road inspections. Among them, intelligent detection of road cracks based on deep learning has become one of the research hotspots in this field.
[0003] Before the wide application of deep learning technology, the intelligent detection of road cracks mainly relied on traditional two-dimensional image analysis and processing systems. The method framework can be summarized into typical paradigms such as feature extraction (including gradient feature extraction, gray threshold segmentation), texture feature analysis, frequency domain analysis, and minimum path search algorithms. Although such algorithms are applicable in specific scenarios (such as the identification of linear cracks in a simple background), they generally have technical defects such as large computational amounts, sensitivity to environmental factors such as light and weather changes, and insufficient robustness in practical engineering applications. In addition, since the shape of road cracks is not regular, the assumption models used in traditional image processing methods often deviate significantly from the morphological characteristics of real cracks, which fundamentally restricts the engineering application value of traditional algorithms.
[0004] In the process of breaking through the bottleneck of traditional detection methods, the rapid evolution of artificial intelligence technology has prompted researchers to introduce diverse machine learning solutions for road damage identification, such as typical integrated algorithms like Support Vector Machine (SVM), Random Forest, Markov Random Field, and Adaboost. The paradigm innovation of deep learning has promoted the revolutionary application of Convolutional Deep Neural Network (DCNN) in road damage detection. With its end-to-end feature learning ability, it has become the core breakthrough direction in this field. Compared with traditional image segmentation methods that rely on manual parameter adjustment and threshold setting, DCNN realizes the automation of the parameter optimization process through an adaptive feature extraction mechanism with a large amount of labeled data, gets rid of the prior restrictions on the geometric features of cracks, and shows better environmental adaptability in complex scenarios such as illumination changes and stain interference. From the dimension of technical implementation, the data-driven road damage detection system mainly covers three sub-architectures: an image classification model based on global feature discrimination, an object detection framework focusing on local damage localization, and a semantic segmentation network for pixel-level analysis. Among them, the image classification model focuses on the macroscopic discrimination between sound road surfaces and damaged road surfaces; the object detection system realizes the instance recognition of multi-class cracks through bounding boxes; while the semantic segmentation technology accurately depicts the morphological features of cracks through pixel-level mask generation.
[0005] Although existing intelligent road crack detection algorithms based on deep learning, especially semantic segmentation algorithms, have shown excellent performance, all publicly available algorithms in the field cannot achieve a balance among detection performance, model scale, and model processing speed, which limits the further deployment of these algorithms to road automatic inspection vehicles / robots. Against the background of the accelerating digitalization process of smart cities, researchers need to face real challenges, explore the optimization principles and methods of algorithms for the problems encountered in the practical application of intelligent road damage recognition algorithms, and develop intelligent road damage detection algorithms more suitable for the deployment of automatic inspection vehicles and robots, so as to accelerate the application pace of these algorithms in the fields of digital twin and smart city construction.
[0006] After retrieval, the Chinese invention patent application publication number CN116758507A discloses a pavement quality analysis method, device and program based on disease image acquisition and segmentation, including the following steps: The vehicle starts; The vehicle wheels drive the wheel speed encoder to trigger the industrial camera to collect road surface image data. At the same time, the satellite positioning and navigation system obtains the longitude and latitude positioning information of the vehicle, and then corresponds the road surface image data with the positioning information one by one; The industrial computer in the vehicle uses the large convolutional U-shaped network structure of the pavement disease segmentation model to perform real-time detection of pavement diseases on the road surface image; Use a three-stage training method to train the pavement disease segmentation model; Use the reparameterization technology to deploy and infer the large convolutional kernel U-shaped network structure to obtain the overall pavement defect detection result; Calculate the pavement condition index PCI according to the detection result; Automatically generate a pavement analysis report and provide pavement maintenance suggestions. This existing patent application has problems such as a relatively large model scale, too high requirements for the hardware configuration of the inspection vehicle, and a relatively slow detection speed.
[0007] How to achieve real-time and accurate detection of road surface cracks based on a lightweight model has become a technical problem to be solved. Summary of the Invention
[0008] The purpose of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a lightweight real-time road surface crack intelligent detection method.
[0009] The purpose of the present invention can be achieved by the following technical solutions:
[0010] According to one aspect of the present invention, a lightweight real-time road surface crack intelligent detection method is provided, and the method includes the following steps:
[0011] Step 1: Standardize the obtained road surface image and convert it into a tensor;
[0012] Step 2: In the training stage, use the training set data to train the lightweight generator network. The lightweight generator network includes a U-shaped main segmentation network and an adaptive boundary extraction module. The U-shaped main segmentation network includes a first encoder for extracting local features and improving the network inference speed, a fast vision transformer for extracting global features, and a first decoder; The adaptive boundary extraction module is used to dynamically capture the boundary features of the training set data;
[0013] Step 3: In the training stage, input the output result of the lightweight generator network into the self-attention discriminator network, train the self-attention discriminator network from both the global and local levels to supervise the output result of the lightweight generation network, and save the model parameters with the best effect on the validation set;
[0014] Step 4: Input the result processed in Step 1 into the trained U-shaped main segmentation network, load the saved model parameters, and output the pixel-level road crack detection result in real time;
[0015] The first encoder includes two residual convolutional double-layer downsampling modules and two convolutional double-layer downsampling modules, and the first decoder includes two residual convolutional double-layer upsampling modules and two convolutional double-layer upsampling modules.
[0016] Preferably, each of the residual convolutional double-layer downsampling modules includes a first max pooling layer, a first depthwise separable convolutional layer, a first batch normalization operation layer, a first ReLU activation layer, a second depthwise separable convolutional layer, a second batch normalization operation layer, and a second ReLU activation layer, wherein the output of the first depthwise separable convolutional layer is cross-connected to the output of the second depthwise separable convolutional layer;
[0017] The structure of each of the convolutional double-layer downsampling modules is similar to that of the residual convolutional double-layer downsampling module, and the convolutional layer therein uses a standard convolutional layer.
[0018] Preferably, each of the residual convolutional double-layer upsampling modules includes a first bilinear upsampling layer, a third depthwise separable convolutional layer, a fifth batch normalization operation layer, a fifth ReLU activation layer, a fourth depthwise separable convolutional layer, a sixth batch normalization operation layer, and a sixth ReLU activation layer, wherein the output of the third depthwise separable convolutional layer is cross-connected to the output of the fourth depthwise separable convolutional layer;
[0019] The structure of each of the convolutional double-layer upsampling modules is similar to that of the residual convolutional double-layer upsampling module, and the convolutional layer therein uses a standard convolutional layer.
[0020] Preferably, the fast vision transformer includes a spatial attention layer, a standard convolutional layer, a spatial attention layer, a channel attention layer, a standard convolutional layer, and a dropout layer connected in sequence.
[0021] Preferably, the adaptive boundary extraction module includes a boundary extraction network sub-module, and the boundary extraction network sub-module is trained to make the extracted boundary information consistent with the boundary ground truth Figure 1 consistent.
[0022] Preferably, the adaptive boundary extraction module uses an adaptive erosion and dilation iterative algorithm to obtain the boundary ground truth map, and the specific method includes:
[0023] First, extract the first-channel image matrix of the input road surface image tensor data and perform data type normalization processing;
[0024] If the kernel radius is not initialized, perform adaptive erosion iteration to calculate the adaptive radius value;
[0025] Calculate the kernel radius using an empirical coefficient based on an adaptive radius value :
[0026] ,
[0027] where round(x) is a rounding function that returns the integer closest to x; max(a, b) is the larger value of a and b; is the adaptive radius value;
[0028] Generate a diamond kernel based on the kernel radius and the Manhattan distance principle, then perform a dilation operation on the original image matrix using this diamond kernel, and perform a difference operation between the dilation result and the original image matrix to obtain the boundary ground truth map corresponding to each image.
[0029] More preferably, if the kernel radius has been initialized, directly generate a diamond kernel based on the final kernel radius and the Manhattan distance principle, then perform a dilation operation on the original image matrix using this diamond kernel, and perform a difference operation between the dilation result and the original image matrix to obtain the boundary ground truth map corresponding to each image.
[0030] More preferably, the process of performing adaptive erosion iteration to calculate the adaptive radius value includes:
[0031] Initialization process: Initialize the current image matrix as a copy of the original image, set the initial value of the adaptive radius to 0, and create an elliptical erosion kernel;
[0032] Erosion iteration operation: Perform iterative erosion operation on the image matrix using the elliptical erosion kernel, and the adaptive radius value increases by 1 after each round of iteration until the image matrix becomes all zeros and the iteration terminates.
[0033] Preferably, the self-attention discriminator network is a U-shaped structure, including a second encoder and a second decoder, and transmits information through skip connections;
[0034] The second encoder consists of four self-attention convolutional downsampling modules, and each self-attention convolutional downsampling module includes a high-efficiency multi-scale self-attention layer, a second batch normalization layer, and a tenth ReLU activation layer connected in sequence;
[0035] The second decoder consists of four convolutional downsampling modules, and each convolutional downsampling module includes a transposed convolution layer, a third batch normalization layer, and an eleventh ReLU activation layer.
[0036] More preferably, the input of the self-attention discriminator network includes a pair of fake images composed of the output of the lightweight generator network and the input road surface image, and a pair of real images composed of the image ground truth and the input road surface image. The self-attention discriminator network is trained through the discriminator local loss and the discriminator global loss, and the per-pixel local consistency and the image global consistency are constrained through adversarial training with the lightweight generator network;
[0037] The efficient multi-scale self-attention layer is based on cross-space learning, reshapes part of the channel dimension into the batch dimension, constructs a parallel network, designs local cross-channel interactions, and uses the cross-space learning method to fuse the output feature maps of the parallel network.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] (1) The lightweight generator network of the present invention extracts global features and local features through the U-shaped main segmentation network, dynamically captures boundary features through the adaptive boundary extraction module, and combines the lightweight generator network to supervise the output results of the lightweight generation network from both global and local levels. It has the characteristics of lightweight, high performance, and high real-time performance, effectively overcoming the difficulty of the prior art in balancing detection performance, model scale, and model processing speed, making it suitable for in-vehicle deployment of intelligent inspection vehicles, and realizing real-time and accurate fully automated intelligent road crack detection.
[0040] (2) The U-shaped main segmentation network of the present invention includes a residual convolution double-layer downsampling module and a residual convolution double-layer upsampling module. While extracting local features of the input image based on the convolutional neural network, the double-layer sampling further improves the inference speed of the network; while the fast vision transformer extracts global features of the input image, it can also improve the inference speed of the network. This lightweight network further improves the real-time performance and detection performance of road cracks.
[0041] (3) The adaptive boundary extraction module of the present invention realizes the dynamic capture of crack contour features through the adaptive erosion and dilation iterative algorithm and the boundary extraction network sub-module, enhances the learning ability of the U-shaped main segmentation network for crack boundary features, can effectively cope with cracks of different shapes and thicknesses, and improves the accuracy and robustness of network detection.
[0042] (4) The present invention designs a self-attention discriminator network, uses a self-attention convolution downsampling module to enhance the feature extraction ability of the input image pair, so as to enhance the discrimination ability of the discriminator network, and designs a local discrimination loss and a global discrimination loss, realizing high-dimensional supervision of the lightweight generator network composed of the U-shaped main segmentation network and the adaptive boundary extraction module, and further improving the crack detection performance of the main segmentation network.
[0043] (5) Experiments were conducted using the publicly available datasets UDTIRI-Crack and Deepcrack. The experimental results show that, compared with the existing publicly available crack detection algorithms, the proposed method achieves the best performance in terms of both detection performance and inference speed. Description of the Drawings
[0044] Figure 1 It is a schematic flowchart of the intelligent road crack detection method in the present invention;
[0045] Figure 2 It is a schematic flowchart of the adaptive boundary extraction algorithm in the present invention;
[0046] Figure 3 It is an example of the road crack image used during the test of the intelligent road crack detection method in the present invention;
[0047] Figure 4 It is an example of the road crack segmentation result during the test of the intelligent road crack detection method in the present invention;
[0048] Figure 5 It is a schematic flowchart of the intelligent road crack detection method in the present invention. Specific Embodiments
[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0050] The objective of the present invention is to overcome the defects existing in the above-mentioned existing deep learning-based road crack detection algorithms (inability to balance detection performance, model scale, and model processing speed), and propose a lightweight, efficient, and novel real-time intelligent road crack detection algorithm suitable for in-vehicle deployment.
[0051] Embodiment 1
[0052] This embodiment relates to a lightweight real-time road crack intelligent detection method, the core of which is to design a lightweight generator network and a self-attention discriminator network. Among them, the lightweight generator network includes a U-shaped main segmentation network and an adaptive boundary extraction module. The U-shaped main segmentation network includes a novel residual convolution double-layer downsampling module and a residual convolution double-layer upsampling module, which are used to extract local features and improve the network inference speed. The adaptive boundary extraction module uses an adaptive erosion and dilation iterative algorithm to dynamically capture the crack contour features. The self-attention discriminator network contains a novel self-attention convolution downsampling module, which is used to enhance the network's feature extraction ability for input image pairs. The lightweight generator network is trained through an adversarial training strategy, and finally a road crack segmentation network with high detection performance, lightweight scale, and fast inference speed is obtained.
[0053] As Figure 1 and Figure 5 , the method includes the following steps:
[0054] Step S1: Standardize the training data and convert it into a tensor as the input of the U-shaped main segmentation network;
[0055] Step S2: In the training stage, a brand-new U-shaped main segmentation network is designed, and a novel residual convolution double-layer downsampling module and a residual convolution double-layer upsampling module are proposed. While extracting the local features of the input image based on the convolutional neural network, the inference speed of the network is improved; at the same time, a novel fast vision transformer is proposed, and while extracting the global features of the input image based on the vision transformer, the inference speed of the network is improved;
[0056] Step S3: In the training stage, an adaptive boundary extraction module is designed, and an adaptive erosion and dilation iterative algorithm and a novel boundary extraction network sub-module are proposed to dynamically capture the crack contour features (i.e., boundary features), enhance the learning ability of the segmentation model for crack boundary features, and can effectively handle cracks of different shapes and thicknesses, achieve better detection results for the diverse shapes and thicknesses of cracks, and improve the robustness of crack detection.
[0057] As Figure 2 shown, the input of the adaptive boundary extraction module is the ground truth of the training set. After the ground truth is preprocessed, the dynamic structure element parameters are calculated by dilation, and the kernel radius is obtained according to the empirical coefficient. According to the obtained kernel radius, a rhombic kernel is constructed, and the input ground truth image is dilated using the rhombic kernel, and then the result of the dilation operation is subtracted from the input image to extract the boundary features and reconstruct the output boundary ground truth map. At the same time, a novel boundary extraction network sub-module is designed inside the model, and the boundary information extracted by the novel boundary extraction network sub-module is trained to be constrained with the boundary ground truth Figure 1So as to enhance the model's learning ability for crack boundary features.
[0058] Step S4: Designed a self-attention discriminator network, proposed a new self-attention convolutional downsampling module to enhance the feature extraction ability of the input image pair, so as to enhance the discrimination ability of the discriminator network. Based on the adversarial training strategy, by constraining the pixel-by-pixel local consistency and the global image consistency, high-dimensional supervision of the generator network composed of the U-shaped main segmentation network and the adaptive boundary extraction module is realized from two aspects of local and global supervision, further improving the crack detection performance of the main segmentation network.
[0059] Specifically, the output of the main segmentation network and the annotation corresponding to the input image are respectively combined with the input image to form image pairs, which are used as the input of the self-attention discriminator network. The designed discriminator network also adopts a U-shaped structure. Its encoder part consists of four self-attention convolutional downsampling modules, and the decoder part consists of four convolutional upsampling modules.
[0060] Step S5: Train the U-shaped main segmentation network, the adaptive boundary extraction module and the self-attention discriminator network. Use the validation set for verification during the training process, and save the model parameters with the best effect on the validation set for use in the test stage, including the weights and biases of each neuron in each layer.
[0061] Step S6: In the test stage, input the road surface crack image and only use the trained U-shaped main segmentation network to obtain the pixel-level road surface crack detection result.
[0062] In the construction of a smart city oriented to digital twins, the method proposed in this invention application can be used as a new intelligent detection tool for road surface damage. Its characteristics of lightweight, high performance, and high real-time make it suitable for in-vehicle deployment on intelligent inspection vehicles, which helps to accumulate background reports on road surface damage and realize intelligent data sharing among systems such as road construction, management, and maintenance. Compared with the existing deep learning-based road surface crack detection algorithms, this algorithm achieves a balance among detection performance, model scale, and model processing speed.
[0063] In step S1, in the pre-training stage, the training data is standardized and converted into tensors as the input for the subsequent network; the generator network is trained to learn the mapping from the input road surface crack map to the crack segmentation result. For the Deepcrack dataset, since it has no validation set, the original training set is randomly divided into a training set and a validation set at a ratio of 7:3. For example, for Deepcrack, if the training data is 300 images, then 210 of them are randomly selected as the training set, and the remaining 90 are used as the validation set.
[0064] In step S2, the main segmentation network is a U-shaped network built based on the PyTorch framework, which includes a first encoder and a first decoder. The first encoder includes: two residual convolutional double-layer downsampling modules and two convolutional double-layer downsampling modules.
[0065] Each residual convolutional double-layer downsampling module includes a first max pooling layer, a first depthwise separable convolutional layer, a first batch normalization operation layer, a first ReLU activation layer, a second depthwise separable convolutional layer, a second batch normalization operation layer, and a second ReLU activation layer. Among them, a cross-layer connection structure is designed from the output of the first depthwise separable convolutional layer to the output of the second depthwise separable convolutional layer;
[0066] Each convolutional double-layer downsampling module includes a second max pooling layer, a first standard convolutional layer, a third batch normalization operation layer, a third ReLU activation layer, a second standard convolutional layer, a fourth batch normalization operation layer, and a fourth ReLU activation layer.
[0067] The first decoder includes: two residual convolutional double-layer upsampling modules and two convolutional double-layer upsampling modules.
[0068] Each residual convolutional double-layer upsampling module includes a first bilinear upsampling layer, a third depthwise separable convolutional layer, a fifth batch normalization operation layer, a fifth ReLU activation layer, a fourth depthwise separable convolutional layer, a sixth batch normalization operation layer, and a sixth ReLU activation layer. Among them, a cross-layer connection structure is designed from the output of the third depthwise separable convolutional layer to the output of the fourth depthwise separable convolutional layer;
[0069] Each convolutional double-layer upsampling module includes a second bilinear upsampling layer, a third standard convolutional layer, a seventh batch normalization operation layer, a seventh ReLU activation layer, a fourth standard convolutional layer, an eighth batch normalization operation layer, and an eighth ReLU activation layer.
[0070] The first encoder and the first decoder are connected by a fast vision transformer module, which is used to further extract the global features of the input image and improve the inference speed of the network. The fast vision transformer module includes a spatial attention layer, a standard convolutional layer, a spatial attention layer, a channel attention layer, a standard convolutional layer, and a dropout layer connected in sequence.
[0071] Through the designed residual convolutional double-layer downsampling module, residual convolutional double-layer upsampling module, and fast vision transformer, the main segmentation network can extract the global features and local features of the input image, so as to obtain high detection performance and maintain high inference speed.
[0072] The segmentation loss function of the main segmentation network used in this step is defined as follows:
[0073] ,
[0074] Among them, represents the BCE loss function (Binary Cross-Entropy Loss), represents the Dice loss function, and α and β are weight coefficients respectively, used to balance the importance of the two losses. Among them, the BCE loss function is used to improve the accuracy of single-pixel classification, and the Dice loss function is used to optimize the overall overlap between the predicted region and the ground truth label. Using them together can achieve the purpose of improving the robustness and segmentation accuracy of the model.
[0075] Specifically, the formula of the BCE loss function is as follows:
[0076] ,
[0077] Among them, represents the crack dataset; respectively represent the image and the corresponding ground truth in the dataset; S represents the generator network; is the Sigmoid function, which converts the output of the network into a probability value; log is the natural logarithm.
[0078] The Dice loss function is as follows:
[0079] ,
[0080] Among them, represents the crack dataset; represents the image and the corresponding ground truth in the dataset; S represents the generator network; is a smoothing term, used to avoid the denominator being zero.
[0081] In addition, in the training stage, we designed a self-attention discriminator network and used an adversarial training strategy. For the lightweight generator network composed of the main segmentation network and the adaptive boundary extraction module, the generator loss function can be written in the following form:
[0082] ,
[0083] ,
[0084] ,
[0085] Among them, represents the encoder of the self-attention discriminator; represents the discrimination result of the discriminator at the pixel point (i, j); represents the output result of the lightweight generator network; Represents the global loss function of the generator, which is used to constrain the lightweight generator network to output detection results equivalent to the ground truth at the global level; Represents the local loss function of the generator, which is used to constrain the lightweight generator network to output detection results equivalent to the ground truth at the pixel-level local level; γ Represents the weight coefficient; Represents the expectation function.
[0086] In step S2, convolution is a mathematical concept defined as the serial integral operation of a convolution kernel (also called a filter) function on an input signal. Depthwise separable convolution is an optimized convolution operation widely used in lightweight neural networks. It decomposes the standard convolution into two steps: depthwise convolution and pointwise convolution. Depthwise convolution performs convolution operations independently on each input channel without mixing between channels; pointwise convolution then performs a cross-channel linear combination on the results of depthwise convolution through 1x1 convolution. This method significantly reduces the computational amount and the number of parameters while maintaining the model performance. Batch normalization technology is an operation node added before the non-linear activation function of each layer of the neural network. The operation performed is to normalize the input values and then map them to the range we want. The full name of the ReLU activation function is the Linear rectification function, also known as the rectified linear unit, which is a commonly used non-linear activation function in artificial neural networks. This function is completely an identity function when the input is greater than 0, which is equivalent to not performing any calculation. Therefore, compared with the traditional neural network activation function Sigmoid, it has excellent features such as fast calculation and convenient backpropagation of errors. At the same time, since it is divided into two discontinuous parts at the position of 0, it has the same non-linear characteristics as the Sigmoid function, so it is more suitable for deep feedforward neural networks. The spatial attention mechanism aims to enhance the model's ability to focus on important regions in the image. It dynamically adjusts the importance of different positions by weighting the spatial dimensions (width and height) of the feature map. The specific implementation usually extracts channel information through global average pooling and global max pooling, then generates a spatial attention map, and multiplies it with the original feature map to highlight the key regions and suppress the irrelevant regions. The channel attention mechanism focuses on adjusting the importance of different channels in the feature map. It analyzes the global information, learns the correlation between each channel, and generates a channel attention weight vector. This mechanism usually includes two steps: feature squeezing, obtaining a channel descriptor through global pooling operations; feature excitation, using a fully connected layer or a convolutional layer to model the dependencies between channels and generate attention weights. Finally, these weights are used to recalibrate the channel features to improve the model's ability to capture key information. The core component of the Vision Transformer is the self-attention mechanism, which allows the model to focus on the information of all other positions in the sequence when processing each position of the input sequence. This mechanism generates new representations by calculating the correlation (attention weights) between each element in the input sequence and other elements and then performing a weighted sum on the input according to these weights. To capture different features and patterns, the Fast Vision Transformer uses the multi-head attention mechanism. It divides the input into multiple subspaces and independently applies the self-attention mechanism in each subspace. Then, the outputs of these subspaces are concatenated and linearly transformed to obtain the final attention output.This approach enhances the expressiveness of the model.
[0087] In step S3, in order to improve the segmentation performance of the generator, the present application designs an adaptive boundary extraction module so that the model pays more attention to boundary features after training, thereby improving the detection performance. The execution of this module is divided into two steps. First, before starting training, the boundary of the image truth value in the training set is extracted by the designed adaptive corrosion and expansion iterative algorithm to obtain the boundary truth map corresponding to each image in the training set; second, a new boundary extraction network submodule is designed to extract the boundary information of the input image during training. On the basis of these two steps, the present invention designs a boundary-aware loss function, trains the boundary extraction effect of the boundary extraction network submodule, and fuses its output boundary information into the main segmentation network to enhance the detection performance of the main segmentation network.
[0088] In step S3, specifically, the designed adaptive corrosion and expansion iterative algorithm is as follows: Figure 2 As shown, it includes: obtaining the input image tensor data and the kernel radius, extracting the first channel of the input image tensor data (i.e. Figure 2 The first channel in the image matrix and normalize its data type (refer to Figure 2 in to an unsigned 8-bit integer matrix) to meet the requirements of subsequent morphological operations, expressed as .
[0089] 1) If the kernel radius has been initialized (i.e., the kernel radius required for the specified dilation operation is entered When , a diamond kernel is generated based on the Manhattan distance principle, and its principle can be expressed as:
[0090] ,
[0091] in, r is the radius parameter of the rhombus core; x and y They represent the coordinate offset of a point in the matrix with the center of the matrix as the origin. The non-zero points of the generated kernel need to satisfy the requirement that the sum of the absolute values of the offsets in the two axis directions is less than the radius parameter. r .
[0092] 2) If the kernel radius is not initialized (that is, when the kernel radius parameter is not specified in the input), the algorithm performs the designed adaptive erosion iterative calculation, uses the erosion kernel to erode the image matrix, and calculates the kernel radius. First, the current image matrix is initialized to a copy of the original image, and the initial value of the adaptive radius is set to zero. Then, an elliptical structure element with a size of 3×3 is used for iterative erosion. After each round of iteration, the adaptive radius value is Increment by 1 until the image matrix is completely zeroed (i.e. empty) to terminate the iteration. Based on the obtained adaptive radius value, the kernel radius is calculated using the empirical coefficient:
[0093] ,
[0094] Among them, round(x) is a rounding function that returns the integer closest to x; max(a, b) is a function that returns the larger value between a and b.
[0095] After obtaining the nuclear radius , a diamond-shaped nucleus is generated based on the Manhattan distance principle in the same way. After generating the diamond-shaped nucleus, the original image matrix is then dilated using this diamond-shaped nucleus:
[0096] ,
[0097] Among them, D represents the dilation operation, represents the above operation of obtaining the diamond-shaped nucleus according to the nuclear radius, represents the image after the dilation operation. Subsequently, the dilation result is subtracted from the original image matrix to obtain the boundary ground truth map corresponding to each image, which is expressed as:
[0098] .
[0099] In step S3, specifically, the designed new boundary extraction network sub-module includes a deformable convolutional layer, a first batch normalization layer, a ninth ReLU activation layer, and a labeled convolutional layer connected in sequence. During the training process, the boundary prediction result of the input image is obtained by this module of the present invention, and the formula is expressed as:
[0100] ,
[0101] Among them, , represents the feature map of the boundary prediction result, and H and W are the spatial dimensions of the feature map respectively; x is the feature map output by the first residual convolutional double-layer downsampling module in the first encoder, ; Dconv represents a deformable convolution with a size of 3×3, Conv represents a convolutional layer with a size of 1×1, represents being processed by the activation function sigmoid, and it can be expressed as: .
[0102] Next, expand channels to obtain , combine it with the feature map x to obtain a new feature map which is expressed as:
[0103] ,
[0104] Transfer the new feature map back to the first encoder, which can increase the model's perception of the boundary information of the input image.
[0105] Meanwhile, the present invention designs a boundary loss function based on the Dice function, and by constraining the prediction result of the boundary extraction network sub-module to be consistent with the boundary ground truth Figure 1 it improves the model's ability to extract boundary features of the input image, thereby enhancing the detection effect and detection robustness of the main segmentation network. The loss function used is defined as follows:
[0106] ,
[0107] wherein, is the boundary ground truth map of the input image, is the boundary extraction prediction result of the boundary extraction network.
[0108] In step S4, different from the discriminator framework of the classical GAN which is a classification network, we design a U-shaped self-attention discriminator network, which includes two parts: a second encoder and a second decoder, and transmits information through skip connections. The second encoder consists of four new self-attention convolutional downsampling modules. Each new self-attention convolutional downsampling module includes an efficient multi-scale self-attention layer, a second batch normalization layer, and a tenth ReLU activation layer connected in sequence. Compared with the standard convolutional downsampling module, it can effectively enhance the feature extraction ability for the input image pair, thereby enhancing the discrimination ability of the self-attention discriminator network. Among them, the designed efficient multi-scale self-attention layer is based on cross-space learning. By reshaping part of the channel dimension into the batch dimension, it avoids the dimensionality reduction brought by the standard convolutional layer, and constructs a parallel network, designs local cross-channel interactions, and uses the cross-space learning method to fuse the output feature maps of the parallel network.
[0109] The second decoder consists of four convolutional downsampling modules. Each convolutional downsampling module includes a transposed convolutional layer, a third batch normalization layer, and an eleventh ReLU activation layer. The inputs of the self-attention discriminator network are respectively a pair of fake images composed of the output of the lightweight generator network and the input image, and a pair of real images composed of the image ground truth and the input image. And, we propose a discriminator global loss and a discriminator local loss to train the discriminator to distinguish between fake and real image pairs from both the global and local levels. Finally, through the process of adversarial training with the lightweight generator network, it constrains the pixel-by-pixel local consistency and the global image consistency, and supervises the lightweight generator network from a high-dimensional level to output higher-performance detection results.
[0110] The discriminator loss function used is defined as follows, including the discriminator global loss and the discriminator local loss:
[0111] ,
[0112] ,
[0113] Among them, represents the second encoder of the self-attention discriminator network; and respectively represent the determination results of the self-attention discriminator network on the real image pair and the fake image pair at the pixel point (i, j); represents the output result of the lightweight generator network; I represents the input image, represents the expected function of the real image pair composed of the image ground truth Y and the input image I, represents the expected function of the fake image pair composed of the output of the lightweight generator network and the input image I, represents the discriminator global loss, represents the discriminator local loss.
[0114] In step S5, the graphics card used during training is NVIDIA RTX 3090, the video memory is 24GB, the system used for training is the Ubuntu system, the programming language and version used are Python3.8, and the deep learning framework used is Pytorch 1.11.0.
[0115] In step S5, a total of 200 epochs were trained. Among them, we added the adversary starting from the 50th round. We finally retained the model parameters with the best performance on the validation set. The model parameters of the 50th round, 100th round, 150th round, and 200th round, including the weights and biases of each neuron in each layer, were saved using the torch.save() function in the pytorch framework. During the training process, we used AdamW (Adaptive Moment Estimation with Weight Decay) to optimize the generator, with a momentum value of 0.9 and a weight decay set to 10 -4 . The initial learning rate was set to 0.001, and this learning rate was exponentially decayed in each round with a decay coefficient of 0.99. By the 50th round, we added the discriminator and needed to adjust the learning rate to start exponentially decaying from 0.0001 with a decay coefficient of 0.99. For the discriminator, we used Adam (Adaptive Moment Estimation) for optimization. Among them, the first-order momentum parameter of the discriminator optimizer was set to 0.9, and the second-order momentum parameter was set to 0.99.
[0116] In step S6, tests are conducted on two datasets respectively, including UDTIRI-Crack and Deepcrack. Among them, UDTIRI-Crack contains 2,500 high-quality pavement crack images with a size of 320×320 pixels. These images are collected based on seven public datasets, namely Crack500, CrackLS315, CrackSC, CrackTree260, CRKWH100, DeepCrack537, and ShadowCrack. After being cropped, scaled, etc., they are divided into 1,500 training data images, 400 validation data images, and 600 test data images. Since the UDTIRI-Crack dataset is sourced from different public datasets, it has various characteristics such as crack types, crack morphologies, thicknesses, etc., and is affected by different degrees of shadows, occlusions, lighting conditions, noise, etc.
[0117] In step S6, the DeepCrack dataset contains 537 concrete surface images with a size of 544×384 pixels, having cracks with multi-scales and multi-scenes. All images are manually annotated with pixel-level labels and are divided into two main subsets: 300 images for training and 237 images for testing.
[0118] In step S6, the pavement crack images in the test set are input into the main segmentation network, and the pre-trained model parameters are loaded. The model can then output pixel-level crack detection results. The quantitative evaluation metrics used include precision, recall, accuracy, IoU, F1-Score, and AIoU.
[0119] Compared with the existing deep learning-based pavement crack detection algorithms, this method achieves a balance among detection performance, model scale, and model processing speed. After establishing a complete and scientific pavement damage database, data intelligent sharing can be realized among systems such as pavement construction, management, and maintenance, and a pavement damage prediction model can be further constructed to achieve the "predictable" evolution of pavement damage. In this way, a more reasonable maintenance plan can be formulated, maintenance measures can be predicted and implemented in advance, and thus the maintenance cycle can be reduced.
[0120] Example 2
[0121] This example relates to a lightweight real-time pavement crack intelligent detection method. The lightweight, high-performance, and high-real-time characteristics of this method make it suitable for on-vehicle deployment of intelligent inspection vehicles, contribute to fully automated intelligent detection, and are more suitable for the construction of smart cities in the field of digital twins.
[0122] In the training stage, the present invention includes a designed U-shaped main segmentation network and an adaptive boundary extraction module to form a lightweight generator network, as well as a novel self-attention discriminator network. The U-shaped main segmentation network is trained to learn the mapping from the input image to the pixel-level crack detection result. Among them, we propose a novel residual convolutional double-layer downsampling module and a residual convolutional double-layer upsampling module, which improve the inference speed of the network while extracting the local features of the input image based on the convolutional neural network; propose a novel fast vision transformer, which improves the inference speed of the network while extracting the global features of the input image based on the vision transformer; propose an adaptive erosion and dilation iterative algorithm and a novel boundary extraction network sub-module, which realizes the dynamic capture of the crack contour features, enhances the learning ability of the designed U-shaped main segmentation network for crack boundary features, can effectively handle cracks of different shapes and thicknesses, and improves the robustness of network detection; propose a novel self-attention convolutional downsampling module to enhance the feature extraction ability of the input image pair, so as to enhance the discrimination ability of the discriminator network, and propose a local discrimination loss and a global discrimination loss, which realizes the high-dimensional supervision of the generator network composed of the U-shaped main segmentation network and the adaptive boundary extraction module, and further improves the crack detection performance of the main segmentation network.
[0123] In the testing stage, the model parameters with the best performance on the saved validation set are loaded, the pavement crack image is input, and the pre-trained main segmentation network model is used to obtain the pixel-level crack detection result.
[0124] Figure 3 、 Figure 4 Examples of the pavement crack image and the corresponding segmentation result used in the algorithm test are shown respectively.
[0125] The invention conducts experiments on the UDTIRI-Crack and Deepcrack datasets. The quantitative experimental results of performance comparison are shown in Table 1 and Table 2 respectively, and the quantitative experimental results of the number of model parameters and running speed are shown in Table 3.
[0126] Table 1
[0127] Method Precision (%) Recall (%) Accuracy (%) F1-Score (%) IoU (%) AIoU (%) Deep-Convolutional Crack Detection Network (Deepcrack18) 74.736 58.702 98.457 65.756 48.982 73.708 Deep Hierarchical-Feature-Learning Crack Detection Network (Deepcrack19) 72.831 57.795 98.390 64.448 47.544 72.956 Skip Connected Crack Detection Network (SCCDNet) 71.294 58.371 98.356 64.189 47.263 72.797 Crack-Attention-Network (Crack-Att) 68.790 66.447 98.392 67.598 51.055 74.710 Crack-Dense-Link-Network (CDLN) 56.947 82.918 97.986 67.522 50.968 74.456 Locally Enhanced Cross-shaped Windows Transformer (LECSFormer) 74.711 65.341 98.567 69.712 53.506 76.025 Convolutional-Transformer Crack Segmentation Network (CT-crackseg) 75.019 66.694 98.599 70.612 54.573 76.574 The present invention 73.435 71.225 98.623 72.313 56.633 77.616
[0128] Table 2
[0129] Method Precision (%) Recall (%) Accuracy (%) F1-Score (%) IoU (%) AIoU (%) Deep-Convolutional Crack Detection Network (Deepcrack18) 91.793 68.036 98.354 78.149 64.135 81.219 Deep Hierarchical-Feature-Learning Crack Detection Network (Deepcrack19) 90.300 73.891 98.527 81.276 68.458 83.468 Skip Connected Crack Detection Network (SCCDNet) 70.567 82.726 97.760 76.164 61.504 79.590 Crack-Attention-Network (Crack-Att) 86.868 71.090 98.284 78.191 64.191 81.210 Crack-Dense-Link-Network (CDLN) 78.939 85.833 98.396 82.242 69.839 84.087 Locally Enhanced Cross-shaped Windows Transformer (LECSFormer) 85.935 79.124 98.536 82.389 70.052 84.268 Convolutional-Transformer Crack Segmentation Network (CT-crackseg) 93.733 71.043 98.541 80.825 67.821 83.158 The present invention 89.081 78.795 98.665 83.623 71.856 85.237
[0130] Table 3
[0131] Method Number of parameters (M) FPS Deep-Convolutional Crack Detection Network (Deepcrack18) 30.905 43.935 Deep Hierarchical-Feature-Learning Crack Detection Network (Deepcrack19) 14.72 276.695 Skip Connected Crack Detection Network (SCCDNet) 31.705 113.317 Crack-Attention-Network (Crack-Att) 45.804 58.257 Crack-Dense-Link-Network (CDLN) 19.151 45.934 Locally Enhanced Cross-shaped Windows Transformer (LECSFormer) 16.528 65.46 Convolutional-Transformer Crack Segmentation Network (CT-crackseg) 22.882 24.926 本发明 3.856 338.64
[0132] The experimental results show that: compared with other publicly available road crack detection algorithms in the field, the method of the present invention has achieved leading results in terms of segmentation performance (precision, recall, accuracy, F1-score, IOU, AIoU) and inference speed (FPS) on these two datasets.
[0133] Embodiment 3
[0134] The electronic device of the present invention includes a central processing unit (CPU), which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or computer program instructions loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The CPU, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.
[0135] Multiple components in the device are connected to the I / O interface, including: an input unit, such as a keyboard, a mouse, etc.; an output unit, such as various types of displays, speakers, etc.; a storage unit, such as a magnetic disk, an optical disc, etc.; and a communication unit, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit allows the device to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0136] The processing unit executes the various methods and processes described above. For example, in some embodiments, the method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more steps of the method described above can be executed. Alternatively, in other embodiments, the CPU can be configured to execute the method by any other suitable means (e.g., by means of firmware).
[0137] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.
[0138] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0139] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0140] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A lightweight real-time intelligent road crack detection method, characterized in that: The method comprises the following steps: Step 1: Standardize the acquired road surface image and convert it into a tensor; Step 2: training a lightweight generator network using training set data, wherein the lightweight generator network includes a U-shaped main segmentation network and an adaptive boundary extraction module, wherein the U-shaped main segmentation network includes a first encoder for extracting local features and improving network inference speed, a fast visual transformer for extracting global features, and a first decoder; the adaptive boundary extraction module is used to dynamically capture boundary features of the training set data; Step 3: Input the output of the lightweight generator network into the self-attention discriminator network, train the self-attention discriminator network from both global and local levels to supervise the output of the lightweight generator network, and save the model parameters with the best effect on the validation set; Step 4: Input the processed result of step 1 into the trained U-shaped main segmentation network, load the saved model parameters, and output the pixel-level pavement crack detection result in real time; The first encoder includes two residual convolution double-layer down-sampling modules and two convolution double-layer down-sampling modules, and the first decoder includes two residual convolution double-layer up-sampling modules and two convolution double-layer up-sampling modules; The fast visual transformer comprises a spatial attention layer, a standard convolution layer, a spatial attention layer, a channel attention layer, a standard convolution layer and a random inactivation layer connected in sequence; The adaptive boundary extraction module includes a boundary extraction network submodule, and the boundary extraction network submodule is trained and constrained so that the extracted boundary information is consistent with the boundary truth map.
2. A lightweight real-time intelligent detection method for pavement cracks according to claim 1, characterized in that: Each of the residual convolutional double-layer downsampling modules includes a first maximum pooling layer, a first depth-separable convolutional layer, a first batch of regularization operation layers, a first ReLU activation layer, a second depth-separable convolutional layer, a second batch of regularization operation layers, and a second ReLU activation layer, wherein the output of the first depth-separable convolutional layer is cross-layer connected to the output of the second depth-separable convolutional layer; The structure of each of the convolutional double-layer downsampling modules is similar to that of the residual convolutional double-layer downsampling module, wherein the convolutional layer adopts a standard convolutional layer.
3. A lightweight real-time intelligent detection method for pavement cracks according to claim 1, characterized in that: Each of the residual convolutional double-layer upsampling modules includes a first bilinear upsampling layer, a third depth-separable convolutional layer, a fifth batch regularization operation layer, a fifth ReLU activation layer, a fourth depth-separable convolutional layer, a sixth batch regularization operation layer and a sixth ReLU activation layer, wherein the output of the third depth-separable convolutional layer is cross-layer connected to the output of the fourth depth-separable convolutional layer; The structure of each of the convolutional double-layer upsampling modules is similar to that of the residual convolutional double-layer upsampling module, wherein the convolutional layer adopts a standard convolutional layer.
4. A lightweight real-time intelligent detection method for pavement cracks according to claim 1, characterized in that: The adaptive boundary extraction module adopts an adaptive corrosion and expansion iterative algorithm to obtain the boundary truth map, and the specific method includes: First, the first channel image matrix of the input road surface image tensor data is extracted and the data type is standardized; If the core radius is not initialized, the adaptive erosion iteration is performed to calculate the adaptive radius value; Based on the adaptive radius value, the kernel radius is calculated using an empirical coefficient : , Among them, round(x) is a rounding function that returns the integer closest to x; max(a,b) is to take the larger value of a and b; is the adaptive radius value; A diamond kernel is generated based on the kernel radius and Manhattan distance principle, and then the diamond kernel is used to dilate the original image matrix. The dilation result is differentially calculated with the original image matrix to obtain the boundary truth map corresponding to each image.
5. A lightweight real-time intelligent detection method for pavement cracks according to claim 4, characterized in that: If the kernel radius has been initialized, a diamond kernel is directly generated based on the final kernel radius and the Manhattan distance principle, and then the diamond kernel is used to dilate the original image matrix, and the dilation result is differentially calculated with the original image matrix to obtain the boundary truth map corresponding to each image.
6. A lightweight real-time intelligent road crack detection method according to claim 5, characterized in that: The process of performing adaptive erosion iteration to calculate the adaptive radius value includes: Initialization process: Initialize the current image matrix to be a copy of the original image, set the initial value of the adaptive radius to 0, and create an elliptical erosion kernel; Iterative erosion operation: an elliptical erosion kernel is used to iteratively erode the image matrix, and the radius value is adaptive after each iteration. Increment by 1 until the image matrix is completely zeroed and the iteration is terminated.
7. A lightweight real-time intelligent road crack detection method according to claim 1, characterized in that: The self-attention discriminator network is a U-shaped structure, comprising a second encoder and a second decoder, and transmits information through a jump connection; The second encoder is composed of four self-attention convolutional downsampling modules, each of which includes an efficient multi-scale self-attention layer, a second batch normalization layer, and a tenth ReLU activation layer connected in sequence; The second decoder is composed of four convolutional downsampling modules, each of which includes a deconvolution layer, a third batch normalization layer and an eleventh ReLU activation layer.
8. A lightweight real-time intelligent road crack detection method according to claim 7, characterized in that: The input of the self-attention discriminator network includes a fake image pair consisting of the output of the lightweight generator network and the input road image, and a real image pair consisting of the image truth value and the input road image, the self-attention discriminator network is trained by the discriminator local loss and the discriminator global loss, and the pixel-by-pixel local consistency and the image global consistency are constrained by adversarial training with the lightweight generator network; The efficient multi-scale self-attention layer is based on cross-space learning, which reshapes part of the channel dimensions into batch dimensions, builds a parallel network, designs local cross-channel interactions, and uses a cross-space learning method to fuse the output feature maps of the parallel networks.
Citation Information
Patent Citations
Pavement quality analysis method, device and program based on disease image acquisition and segmentation
CN116758507A
Multi-target multi-size image detection method and system under complex background, electronic equipment and storage medium
CN116229086A
Electric imaging logging image crack segmentation method and system based on deep learning
CN116934780A