A crack detection method based on pixel intensity sequential change and segmentation network
By combining traditional machine learning and deep learning, and using the sequential changes in pixel intensity to construct a crack segmentation network, the problems of low crack detection accuracy and high cost are solved, and efficient and accurate crack detection is achieved, which is suitable for infrastructure such as buildings, roads and bridges.
Patent Information
- Application Number
- CN202310447441.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-24
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2043-04-24
AI Technical Summary
Existing crack detection methods have problems such as low detection accuracy, high cost, and low efficiency. Especially in diverse crack detection scenarios, the relationship between cracks and non-crack areas is complex and the sample distribution is unbalanced, which increases the difficulty of detection.
Combining traditional machine learning and deep learning methods, the sequential changes in pixel intensity are used to construct a crack segmentation network. Through the encoder-decoder architecture, logic gating units and attention mechanism, an auxiliary graph is generated to guide the model to perceive crack features. The weighted binary cross loss function and Dice similarity coefficient are combined to optimize model training.
It improves the accuracy and efficiency of crack detection, reduces detection costs, and achieves efficient, accurate and robust crack detection, which is applicable to a variety of infrastructure.
Smart Images

Figure CN116452564B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer vision, and in particular to a crack detection method based on pixel intensity sequential change and segmentation network. BACKGROUND
[0002] When we talk about the safety and stability of infrastructure, cracks are a ubiquitous problem. This structural defect has a serious impact on the strength and stability of infrastructure, which can lead to adverse consequences such as building collapse, road collapse, bridge rupture, etc., thereby endangering people's life and property safety and social and economic development. Traditional crack solutions include manual detection and repair; manual detection requires a lot of manpower and time, and accuracy and consistency are affected by personnel subjective factors; repair methods are usually based on experience and professional knowledge, lack of scientific nature and precision, and are costly and have a long maintenance cycle, which can cause infrastructure downtime and losses.
[0003] To solve these problems, we need more intelligent and efficient methods to deal with cracks. Today, intelligent solutions mainly use traditional machine learning algorithms and deep learning techniques. Traditional machine learning algorithms are suitable for small sample datasets, have strong stability and cost-effectiveness, but only model locally, lack a global view, and do not perform well in terms of accuracy. Deep learning algorithms can learn higher-level abstract features, have superior accuracy, and can learn features from data autonomously, reducing the need for human intervention and improving detection efficiency, but the cost is relatively high, limited by the size and diversity of the data.
[0004] In the context of diversified crack detection, there is a complex relationship between cracks and non-crack areas, which poses a great challenge to the accuracy of detection methods. In addition, due to the relatively small number of crack samples, the imbalance of sample distribution also makes crack detection more difficult. Therefore, in practical applications, crack detection methods need to have characteristics such as high efficiency, accuracy, and robustness. SUMMARY
[0005] The present application solves the problems existing in the prior art crack detection method and provides a crack detection method based on pixel intensity sequential change and segmentation network, which has the characteristics of high efficiency, accuracy, and robustness.
[0006] To solve the above problems, the present application is implemented by the following technical solutions:
[0007] A crack detection method based on pixel intensity sequential change and segmentation network, comprising the following steps:
[0008] Step 1, construct a crack segmentation network;
[0009] Step 2: Obtain a sample crack data set, and process each sample crack image in the crack data set using the characteristic of sequential changes in pixel intensity to generate an auxiliary image of each sample crack image;
[0010] Step 3: Use the sample crack image and its auxiliary image to train the crack segmentation network in step 1 to obtain a crack segmentation network model;
[0011] Step 4: Acquire the crack image to be detected, and process the crack image to be detected using the characteristic of sequential changes in pixel intensity to generate an auxiliary image of the crack image to be detected;
[0012] Step 5: The crack image to be detected and its auxiliary image are fed into the crack segmentation network model of step 3 to complete crack detection of the crack image to be detected.
[0013] The crack segmentation network in step 1 above consists of 1 conversion layer, 26 self-attention convolutional layers, 3 mixed pooling layers, 2 maximum pooling layers, 2 maximum inverse pooling layers, 3 alternating convolutional layers, 5 scaled attention layers, 1 splicing layer and 1 convolutional layer;
[0014] The input of the conversion layer forms the input of the crack segmentation network; the output of the conversion layer is connected to the input of the first self-attention convolutional layer, the output of the first self-attention convolutional layer is connected to the input of the second self-attention convolutional layer, and the output of the second self-attention convolutional layer is connected to the input of the first mixed pooling layer; the output of the first mixed pooling layer is connected to the input of the third self-attention convolutional layer, the output of the third self-attention convolutional layer is connected to the input of the fourth self-attention convolutional layer, and the output of the fourth self-attention convolutional layer is connected to the input of the second mixed pooling layer; the output of the second mixed pooling layer is connected to the input of the fifth self-attention convolutional layer, the output of the fifth self-attention convolutional layer is connected to the input of the sixth self-attention convolutional layer, and the output of the sixth self-attention convolutional layer is connected to the seventh self-attention convolutional layer The output of the seventh self-attention convolution layer is connected to the input of the third mixed pooling layer; the output of the third mixed pooling layer is connected to the input of the eighth self-attention convolution layer, the output of the eighth self-attention convolution layer is connected to the input of the ninth self-attention convolution layer, the output of the ninth self-attention convolution layer is connected to the input of the tenth self-attention convolution layer, and the output of the tenth self-attention convolution layer is connected to the input of the first maximum pooling layer; the output of the first maximum pooling layer is connected to the input of the eleventh self-attention convolution layer, the output of the eleventh self-attention convolution layer is connected to the input of the twelfth self-attention convolution layer, the output of the twelfth self-attention convolution layer is connected to the input of the thirteenth self-attention convolution layer, and the output of the thirteenth self-attention convolution layer is connected to the input of the second maximum pooling layer;
[0015] The output of the second max-pooling layer is connected to the input of the first max-unpooling layer; the output of the first max-unpooling layer is connected to the input of the fourteenth self-attention convolutional layer, the output of the fourteenth self-attention convolutional layer is connected to the input of the fifteenth self-attention convolutional layer, the output of the fifteenth self-attention convolutional layer is connected to the input of the sixteenth self-attention convolutional layer, the output of the sixteenth self-attention convolutional layer is connected to the input of the second max-unpooling layer; the output of the second max-unpooling layer is connected to the input of the seventeenth self-attention convolutional layer, the output of the seventeenth self-attention convolutional layer is connected to the input of the eighteenth self-attention convolutional layer, the output of the eighteenth self-attention convolutional layer is connected to the input of the nineteenth self-attention convolutional layer, the output of the nineteenth self-attention convolutional layer is connected to the input of the first alternating convolutional layer; the output of the first alternating convolutional layer is connected to the input of the twentieth self-attention convolutional layer, the output of the twentieth self-attention convolutional layer is connected to the input of the twenty-first self-attention convolutional layer, the output of the twenty-first self-attention convolutional layer is connected to the input of the twenty-second self-attention convolutional layer, the output of the twenty-second self-attention convolutional layer is connected to the input of the second alternating convolutional layer; the output of the second alternating convolutional layer is connected to the input of the twenty-third self-attention convolutional layer, the output of the twenty-third self-attention convolutional layer is connected to the input of the twenty-fourth self-attention convolutional layer, the output of the twenty-fourth self-attention convolutional layer is connected to the input of the third alternating convolutional layer; the output of the third alternating convolutional layer is connected to the input of the twenty-fifth self-attention convolutional layer, the output of the twenty-fifth self-attention convolutional layer is connected to the input of the twenty-sixth self-attention convolutional layer;
[0016] The outputs of the eleventh, twelfth, thirteenth, fourteenth, fifteenth and sixteenth self-attention convolutional layers are simultaneously connected to the first scaling attention layer; the outputs of the eighth, ninth, tenth, seventeenth, eighteenth and nineteenth self-attention convolutional layers are simultaneously connected to the second scaling attention layer; the outputs of the fifth, sixth, seventh, twentieth, twenty-first and twenty-second self-attention convolutional layers are simultaneously connected to the third scaling attention layer; the outputs of the third, fourth, twenty-third and twenty-fourth self-attention convolutional layers are simultaneously connected to the fourth scaling attention layer; the outputs of the first, second, twenty-fifth and twenty-sixth self-attention convolutional layers are simultaneously connected to the fifth scaling attention layer; the output of the conversion layer and the outputs of the first to fifth scaling attention layers are simultaneously connected to the input of the splicing layer, the output of the conversion layer is connected to the input of the convolutional layer, and the output of the convolutional layer forms the output of the crack segmentation network.
[0017] The conversion layer of the crack segmentation network comprises one activation layer, one gating logic layer, two residual layers, one spatial attention layer, one channel attention layer, three multiplication layers and one addition layer; the input of the first residual layer and one input of the first and second multiplication layers together form the input of the crack image of the conversion layer; the input of the activation layer forms the input of the auxiliary image of the crack image of the conversion layer, the output of the activation layer is connected to the input of the gating logic layer, and the output of the gating logic layer is simultaneously connected to the other input of the first residual layer and the first multiplication layer; the output of the first multiplication layer is connected to one input of the second residual layer; the output of the first residual layer is divided into two paths, one path is connected to one input of the third multiplication layer through the spatial attention layer, and the other path is connected to the other input of the third multiplication layer through the channel attention layer; the output of the third multiplication layer is simultaneously connected to the other input of the second residual layer and the second multiplication layer; and the outputs of the second residual layer and the second multiplication layer form the output of the conversion layer.
[0018] The mixed pooling layer of the crack segmentation network comprises one horizontal pooling layer, one vertical pooling layer, one average pooling layer and one addition layer; the inputs of the horizontal pooling layer, the vertical pooling layer and the average pooling layer together form the input of the mixed pooling layer, the inputs of the horizontal pooling layer, the vertical pooling layer and the average pooling layer are simultaneously connected to the input of the addition layer, and the output of the addition layer forms the output of the mixed pooling layer.
[0019] The alternating convolution layer of the crack segmentation network comprises two inverse convolution layers, one convolution layer, three activation layers, one subtraction layer and one addition layer; the input of the first inverse convolution layer forms the input of the alternating convolution layer, the output of the first inverse convolution layer is connected to the input of the first activation layer, the output of the first activation layer is connected to the input of the convolution layer, and the output of the convolution layer is connected to the input of the second activation layer; the input of the first inverse convolution layer and the output of the second activation layer are simultaneously connected to the input of the subtraction layer, the output of the subtraction layer is connected to the input of the second inverse convolution layer, and the output of the second inverse convolution layer is connected to the input of the third activation layer; the output of the first activation layer and the output of the third activation layer are simultaneously connected to the input of the addition layer, and the output of the addition layer forms the output of the alternating convolution layer.
[0020] In steps 2 and 4, the specific process of processing the crack image to generate the auxiliary image of the crack image by using the characteristic of the sequential change of the pixel intensity is as follows:
[0021] Step 1) For each pixel point p on the crack image, the pixel point on the crack image with a distance i from the pixel point p in the s direction is regarded as the adjacent pixel point of the pixel point p
[0022] Step 2) Calculate the pixel intensity relative value of each pixel point p of the crack image and its adjacent pixel point
[0023] Step 3) Use the Sigmoid function to compare each pixel p of the crack image with its adjacent pixels The relative value of pixel intensity Normalize and get each pixel point p of the crack image and its adjacent pixels The normalized pixel intensity relative value
[0024] Step 4) Calculate each pixel point p of the crack image and its adjacent pixels Position strength
[0025] Step 5) Calculate the pixel p of the crack image and its adjacent pixels The normalized pixel intensity relative value and position strength The mean value after multiplication is used to obtain the auxiliary pixel intensity f′(p) of each pixel point p in the crack image;
[0026] Step 6) The auxiliary pixel intensity f′(p) of each pixel point p is calculated based on the pixel intensity of each pixel point p in the crack image, thus forming an auxiliary image of the crack image;
[0027] In the above, s=1, 2, ..., n, where n is the number of set directions, and i=1, 2, ..., m, where m is the number of set Euclidean distances.
[0028] In the above step 2), each pixel point p of the crack image and its adjacent pixel points The relative value of pixel intensity for:
[0029]
[0030] Where f(p) represents the pixel intensity of pixel p, Represents the adjacent pixels of pixel p Pixel intensity.
[0031] In the above step 4), each pixel point p of the crack image and its adjacent pixel points Position strength for:
[0032]
[0033] Among them, μ is the set mean, σ is the set variance, ω is the set intensity weight, and i is the pixel p and its adjacent pixels. The Euclidean distance from .
[0034] Compared with the prior art, the present invention has the following characteristics:
[0035] 1. Considering that in the crack detection scenario, crack pixels only occupy a small part of the entire image, while non-crack areas occupy most of the pixels, using only deep learning methods will cause the model to tend to predict that all pixels in the image are non-crack pixels, resulting in reduced detection accuracy. The present invention combines traditional machine learning methods with deep learning methods. Through sample balancing, the model can pay more attention to crack pixels during training, thereby improving the accuracy of crack detection.
[0036] 2. It takes advantage of the traditional machine learning method to quickly and efficiently capture the crack shape features, and assists in the training of deep learning models, thereby more quickly capturing higher-level semantic feature information and improving the accuracy of crack detection. It is significantly superior to similar detection methods in terms of false detection rate, robustness and detection efficiency.
[0037] 3. The segmentation network adopts an encoder-decoder architecture, combined with logic gating units and an attention mechanism, to convert the auxiliary image into a guidance signal that guides the model to perceive crack features, thereby improving the detection efficiency of the segmentation model. Through a novel sampling strategy and a fusion strategy combined with an attention mechanism, the model's ability to perceive cracks is enhanced, thereby improving the accuracy and stability of crack detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 This is the structure diagram of the segmentation network.
[0039] Figure 2 This is the conversion layer structure diagram.
[0040] Figure 3 Graph showing the input and output instances of the conversion layer.
[0041] Figure 4 This is the structure diagram of the mixed pooling layer.
[0042] Figure 5 This is the structure diagram of the alternating convolutional layer.
[0043] Figure 6 This is a sampling test result chart. DETAILED DESCRIPTION
[0044] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to specific examples.
[0045] A crack detection method based on sequential pixel intensity changes and a segmentation network comprises the following steps:
[0046] Step 1: Construct a crack segmentation network.
[0047] The crack segmentation network based on the encoder-decoder is constructed, the auxiliary graph is converted into a guidance signal for strengthening the crack feature perceived by the model through a converter layer, a hybrid pooling layer is used as a down-sampling strategy of the crack segmentation network, so that the model can eliminate a large amount of crack feature irrelevant area information, an alternating convolution layer is used as an up-sampling strategy of the crack segmentation network, so as to strengthen the ability of the segmentation model to perceive detailed features, and a scaling attention mechanism is introduced to fuse feature information generated in different stages of the encoder-decoder, so as to generate a significant and clear crack boundary graph.
[0048] Referring to Figure 1 The crack segmentation network constructed by the application is composed of 1 converter layer, 26 self-attention convolution layers, 3 hybrid pooling layers, 2 maximum pooling layers, 2 maximum inverse pooling layers, 3 alternating convolution layers, 5 scaling attention layers, 1 splicing layer and 1 convolution layer.
[0049] The input of the converter layer forms the input of the crack segmentation network; the output of the converter layer is connected to the input of the first self-attention convolution layer, the output of the first self-attention convolution layer is connected to the input of the second self-attention convolution layer, and the output of the second self-attention convolution layer is connected to the input of the first hybrid pooling layer; the output of the first hybrid pooling layer is connected to the input of the third self-attention convolution layer, the output of the third self-attention convolution layer is connected to the input of the fourth self-attention convolution layer, and the output of the fourth self-attention convolution layer is connected to the input of the second hybrid pooling layer; the output of the second hybrid pooling layer is connected to the input of the fifth self-attention convolution layer, the output of the fifth self-attention convolution layer is connected to the input of the sixth self-attention convolution layer, the output of the sixth self-attention convolution layer is connected to the input of the seventh self-attention convolution layer, and the output of the seventh self-attention convolution layer is connected to the input of the third hybrid pooling layer; the output of the third hybrid pooling layer is connected to the input of the eighth self-attention convolution layer, the output of the eighth self-attention convolution layer is connected to the input of the ninth self-attention convolution layer, the output of the ninth self-attention convolution layer is connected to the input of the tenth self-attention convolution layer, and the output of the tenth self-attention convolution layer is connected to the input of the first maximum pooling layer; the output of the first maximum pooling layer is connected to the input of the eleventh self-attention convolution layer, the output of the eleventh self-attention convolution layer is connected to the input of the twelfth self-attention convolution layer, the output of the twelfth self-attention convolution layer is connected to the input of the thirteenth self-attention convolution layer, and the output of the thirteenth self-attention convolution layer is connected to the input of the second maximum pooling layer.
[0050] The output of the second maximum pooling layer is connected to the input of the first maximum inverse pooling layer; the output of the first maximum inverse pooling layer is connected to the input of the fourteenth self-attention convolutional layer, the output of the fourteenth self-attention convolutional layer is connected to the input of the fifteenth self-attention convolutional layer, the output of the fifteenth self-attention convolutional layer is connected to the input of the sixteenth self-attention convolutional layer, the output of the sixteenth self-attention convolutional layer is connected to the input of the second maximum inverse pooling layer; the output of the second maximum inverse pooling layer is connected to the input of the seventeenth self-attention convolutional layer, the output of the seventeenth self-attention convolutional layer is connected to the input of the eighteenth self-attention convolutional layer, the output of the eighteenth self-attention convolutional layer is connected to the input of the nineteenth self-attention convolutional layer, the output of the nineteenth self-attention convolutional layer is connected to the input of the first alternating convolutional layer; the output of the first alternating convolutional layer is connected to the input of the twentieth self-attention convolutional layer, the output of the twentieth self-attention convolutional layer is connected to the input of the twenty-first self-attention convolutional layer, the output of the twenty-first self-attention convolutional layer is connected to the input of the twenty-second self-attention convolutional layer, the output of the twenty-second self-attention convolutional layer is connected to the input of the second alternating convolutional layer; the output of the second alternating convolutional layer is connected to the input of the twenty-third self-attention convolutional layer, the output of the twenty-third self-attention convolutional layer is connected to the input of the twenty-fourth self-attention convolutional layer, the output of the twenty-fourth self-attention convolutional layer is connected to the input of the third alternating convolutional layer; the output of the third alternating convolutional layer is connected to the input of the twenty-fifth self-attention convolutional layer, the output of the twenty-fifth self-attention convolutional layer is connected to the input of the twenty-sixth self-attention convolutional layer.
[0051] The outputs of the eleventh, twelfth, thirteenth, fourteenth, fifteenth and sixteenth self-attention convolutional layers are simultaneously connected to the first scaling attention layer; the outputs of the eighth, ninth, tenth, seventeenth, eighteenth and nineteenth self-attention convolutional layers are simultaneously connected to the second scaling attention layer; the outputs of the fifth, sixth, seventh, twentieth, twenty-first and twenty-second self-attention convolutional layers are simultaneously connected to the third scaling attention layer; the outputs of the third, fourth, twenty-third and twenty-fourth self-attention convolutional layers are simultaneously connected to the fourth scaling attention layer; the outputs of the first, second, twenty-fifth and twenty-sixth self-attention convolutional layers are simultaneously connected to the fifth scaling attention layer; the output of the conversion layer and the outputs of the first to fifth scaling attention layers are simultaneously connected to the input of the splicing layer, the output of the conversion layer is connected to the input of the convolutional layer, and the output of the convolutional layer forms the output of the crack segmentation network.
[0052] The conversion layer of the above crack segmentation network is mainly to convert the auxiliary graph into a guidance signal for improving the perception of the subsequent segmentation model for cracks. Referring to Figure 2The conversion layer consists of an activation layer, a gated logic layer, two residual layers, a spatial attention layer, a channel attention layer, three multiplication layers, and an addition layer. The first residual layer and one input of the first and second multiplication layers together form the input of the crack image of the conversion layer; the input of the activation layer forms the input of the auxiliary map of the crack image of the conversion layer; the output of the activation layer is connected to the input of the gated logic layer, and the output of the gated logic layer is simultaneously connected to the first residual layer and the other input of the first multiplication layer; the output of the first multiplication layer is connected to one input of the second residual layer; the output of the first residual layer is divided into two paths, one path is connected to one input of the third multiplication layer via the spatial attention layer, and the other path is connected to the other input of the third multiplication layer via the channel attention layer; the output of the third multiplication layer is simultaneously connected to the other input of the second residual layer and the second multiplication layer; the output of the second residual layer and the second multiplication layer form the output of the conversion layer.
[0053] The specific operation process of the conversion layer is as follows:
[0054] 1) Process the auxiliary image to obtain a feature representation mask for the auxiliary image. The input auxiliary image pixel intensity matrix is mapped to the interval [0, 1] using the Sigmoid function to obtain a normalized pixel intensity matrix for the auxiliary image. A gated logic operation is then performed on the normalized pixel intensity matrix of the auxiliary image, i.e., normalized pixel intensity values less than or equal to 0.5 in the normalized pixel intensity matrix of the auxiliary image are set to 0, and normalized pixel intensity values greater than 0.5 in the normalized pixel intensity matrix of the auxiliary image are set to 1, thereby obtaining a feature representation mask for the auxiliary image.
[0055] 2) Based on the feature representation x of the crack image and the feature representation mask of the auxiliary image, the common feature com and the residual feature res of the crack image and its auxiliary image are extracted, where com = x*mask, res = x*(1-mask).
[0056] 3) Spatial attention SA and channel attention ECA are used to refine the residual feature res, that is, M = SA(res)*ECA(res), to help the model adaptively select the image area of interest, thereby improving the segmentation accuracy of the model.
[0057] 4) Combine the common feature com and the residual feature res to obtain the guidance signal RF = M*x+(1-M)*com to improve the crack perception ability of the subsequent segmentation model.
[0058] Figure 3 This is an example diagram of the input and output of the conversion layer. The crack image in the figure is the crack image input to the conversion layer, the auxiliary image is the auxiliary image of the crack image input to the conversion layer, and the signal feature image is the output result of the conversion layer.
[0059] The up-sampling and down-sampling of the crack segmentation network are mainly completed by a mixed pooling layer and an alternating convolution layer, wherein the mixed pooling layer is used as a down-sampling strategy of the crack segmentation network, and the alternating convolution layer is used as an up-sampling strategy of the crack segmentation network.
[0060] The mixed pooling layer is composed of one horizontal pooling layer, one vertical pooling layer, one average pooling layer and one addition layer, as shown in Figure 4 The inputs of the horizontal pooling layer, the vertical pooling layer and the average pooling layer jointly form the input of the mixed pooling layer, the inputs of the horizontal pooling layer, the vertical pooling layer and the average pooling layer are simultaneously connected to the input of the addition layer, and the output of the addition layer forms the output of the mixed pooling layer. In the down-sampling process, considering that the cracks are generally zigzag strip structures, in addition to the traditional average pooling layer, the vertical pooling and the horizontal pooling structure are additionally introduced to extract the features in the vertical and horizontal directions, so as to better capture the features of the crack curve and eliminate a large amount of region information irrelevant to the crack shape features.
[0061] The alternating convolution layer is composed of two inverse convolution layers, one convolution layer, three activation layers, one subtraction layer and one addition layer, as shown in Figure 5 The input of the first inverse convolution layer forms the input of the alternating convolution layer, the output of the first inverse convolution layer is connected to the input of the first activation layer, the output of the first activation layer is connected to the input of the convolution layer, and the output of the convolution layer is connected to the input of the second activation layer; the input of the first inverse convolution layer and the output of the second activation layer are simultaneously connected to the input of the subtraction layer, the output of the subtraction layer is connected to the input of the second inverse convolution layer, and the output of the second inverse convolution layer is connected to the input of the third activation layer; the output of the first activation layer and the output of the third activation layer are simultaneously connected to the input of the addition layer, and the output of the addition layer forms the output of the alternating convolution layer. In the up-sampling process, a repeated convolution module is constructed using a repeated convolution strategy. When the feature map x0 enters the repeated convolution module, the first up-sampling is performed through the inverse convolution layer, and the result is x1. Then, x1 is further refined through the convolution layer to obtain x2, x2 is subtracted from x0 to obtain an intermediate result, and the inverse convolution layer is used for up-sampling again to obtain the second up-sampling result x3. Finally, the two up-sampling results are added, that is, x0 and x3 are added, and the obtained result is used as the final up-sampling result.
[0062] The fusion of the crack segmentation network mainly uses a scaling attention layer, which connects and combines the feature information generated by the crack segmentation network at different stages to generate a significant and clear crack boundary map.
[0063] The loss function of the crack segmentation network adopts a loss function combined with a weighted binary cross loss function and a Dice similarity coefficient, which is used to solve the imbalance problem of foreground and background naturally existing in the crack detection scene. The weighted binary cross loss function can effectively punish the pixel points predicted by the model, so that the model pays more attention to the pixel points predicted correctly.
[0064] Assume is the weight of the non-crack class (foreground class), which is set to 20; is the weight of the crack class, which is set to 1, p is the model prediction result, y is the real sample, and its definition formula can be expressed as:
[0065]
[0066] The Dice similarity coefficient can better reflect the segmentation effect of the model on the minority class, and its definition formula is:
[0067]
[0068] Wherein, epsilon is a smoothing coefficient, which is set to 1e-6, so as to avoid the problem of denominator being 0, and ensure the stability of the loss function.
[0069] Therefore, the final loss function is:
[0070] l=λb bce l bce +λ Dice l Dice
[0071] Wherein, lambda b bce And lambda Dice Are the weights of the two, which are set to 6 and 1.
[0072] The crack segmentation network of the application adopts the architecture of encoder-decoder, and adds the converter, attention mechanism, hybrid pooling operation and alternating convolution strategy, so that the segmentation network can sharpen the semantic features of the crack more accurately and suppress the non-semantic features, improve the extraction of image features and the ability of perceiving detailed features of the network, so that the detection of the crack is more accurate and reliable.
[0073] Step 2, obtain the sample crack data set, and process each sample crack image in the sample crack data set by using the characteristics of the order change of pixel intensity, to generate an auxiliary image of each sample crack image.
[0074] In this example, the sample crack dataset is the CrackSeg9k dataset, downloaded from the Harvard University dataset website. The CrackSeg9k dataset is the largest and most diverse crack segmentation dataset ever constructed, containing 9,255 images. It combines various small open-source datasets. As a result, the dataset exhibits significant diversity in terms of surface, background, lighting, exposure, crack width, and crack type (linear, branching, webbed, and non-crack).
[0075] When processing each sample crack image in the crack sample dataset using the characteristic of sequential changes in pixel intensity to generate auxiliary images for each sample crack image, the relative difference in pixel intensity between the image pixel and other adjacent pixels within a specific distance is calculated, and the calculated relative difference matrix value is mapped to the [0,1] interval and multiplied by the position intensity operator to obtain a single-channel image of relative pixel intensity in each direction. Finally, all single-channel images are averaged, and the result is used as the auxiliary image for the subsequent segmentation model. The specific process is as follows:
[0076] Step 1) For each pixel point p on the crack image, the pixel points at a distance i from the pixel point p in the s direction on the crack image are considered as the adjacent pixel points of the pixel point p.
[0077] Step 2) Calculate the pixel p of each crack image and its adjacent pixels The relative value of pixel intensity
[0078]
[0079] Among them, f(p) represents the pixel intensity of pixel p, Represents the adjacent pixels of pixel p Pixel intensity;
[0080] Step 3) Use the Sigmoid function to compare each pixel p of the crack image with its adjacent pixels The relative value of pixel intensity Normalize and get each pixel point p of the crack image and its adjacent pixels The normalized pixel intensity relative value
[0081]
[0082] Step 4) Calculate each pixel point p of the crack image and its adjacent pixels Position strength
[0083]
[0084] wherein μ is a set mean value, σ is a set variance, ω is a set intensity weight, and i is a pixel point Euclidean distance from the pixel point p;
[0085] Step 5) calculating the normalized pixel intensity relative value of each pixel point p of the crack image and its adjacent pixel points and the mean value after multiplying the position intensity of each pixel point p of the crack image, to obtain the auxiliary pixel intensity f'(p) of each pixel point p of the crack image;
[0086]
[0087] Step 6) according to the auxiliary pixel intensity f'(p) of each pixel point p calculated from the pixel intensity of each pixel point p of the crack image, an auxiliary image of the crack image is constituted.
[0088] The above, s = 1, 2, …, n, n is a set number of directions, in this embodiment, n = 4, representing the up, down, left and right four directions respectively. i = 1, 2, …, m, m is a set number of Euclidean distances, in this embodiment, m = 8.
[0089] The crack extraction based on the order of image pixel intensity is realized by analyzing the pixel intensity difference between the crack area and the non-crack area. It is not only simple and effective, without complex preprocessing and feature extraction of the image, but also has wide applicability, without being limited by the image type and surface characteristics, and is suitable for various types of infrastructure surfaces, such as walls, pavements and bridges, etc. In addition, the auxiliary image is constructed by using the characteristics of the order change of pixel intensity, which can help the segmentation network to perceive the crack area more accurately.
[0090] Step 3, training the crack segmentation network of step 1 using the sample crack image and its auxiliary image to obtain a crack segmentation network model.
[0091] Step 4, collecting the crack image to be detected, and processing the crack image to be detected by using the characteristics of the order change of pixel intensity to generate the auxiliary image of the crack image to be detected.
[0092] Step 4, the method of generating the auxiliary image of the crack image to be detected by using the characteristics of the order change of pixel intensity is the same as the method of generating the auxiliary image of the sample crack image by using the characteristics of the order change of pixel intensity in step 2.
[0093] Step 5, the crack image to be detected and its auxiliary image are input into the crack segmentation network model of step 3, and crack detection of the crack image to be detected is completed.
[0094] Figure 6 For the sampling test result map, it can be seen from the comparison between the manual annotation and the prediction result that the present application can realize accurate crack detection.
[0095] In summary, the present application combines traditional machine learning algorithms and deep learning techniques, not only effectively solves the challenges faced in diversified crack detection scenarios, but also realizes efficient, accurate and robust crack detection, reduces detection cost and improves detection efficiency, and has a wide range of applications, which can be used for crack detection of buildings, roads, bridges and other infrastructure. In practical application, the present application has great potential and can effectively reduce the safety hazards and economic losses caused by cracks, protect people's life and property safety, and promote the development of social economy.
[0096] It should be noted that although the above embodiments of the present application are illustrative, this is not a limitation of the present application, therefore the present application is not limited to the above specific embodiments. Any other embodiments obtained by those skilled in the art under the inspiration of the present application without departing from the principles of the present application are considered to be within the protection scope of the present application.
Claims
1. A crack detection method based on sequential pixel intensity changes and segmentation network, characterized in that: The steps are as follows: Step 1: Construct a crack segmentation network; The crack segmentation network consists of 1 conversion layer, 26 self-attention convolutional layers, 3 mixed pooling layers, 2 maximum pooling layers, 2 maximum inverse pooling layers, 3 alternating convolutional layers, 5 scaled attention layers, 1 splicing layer and 1 convolutional layer; The conversion layer consists of 1 activation layer, 1 gated logic layer, 2 residual layers, 1 spatial attention layer, 1 channel attention layer, 3 multiplication layers and 1 addition layer; the first residual layer and one input of the first and second multiplication layers together form the input of the crack image of the conversion layer; the input of the activation layer forms the input of the auxiliary map of the crack image of the conversion layer, the output of the activation layer is connected to the input of the gated logic layer, and the output of the gated logic layer is simultaneously connected to the first residual layer and the other input of the first multiplication layer; the output of the first multiplication layer is connected to one input of the second residual layer; the output of the first residual layer is divided into two paths, one path is connected to one input of the third multiplication layer via the spatial attention layer, and the other path is connected to the other input of the third multiplication layer via the channel attention layer; the output of the third multiplication layer is simultaneously connected to the other input of the second residual layer and the second multiplication layer; the output of the second residual layer and the second multiplication layer form the output of the conversion layer; The alternating convolution layer consists of 2 deconvolution layers, 1 convolution layer, 3 activation layers, 1 subtraction layer and 1 addition layer; the input of the first deconvolution layer forms the input of the alternating convolution layer, the output of the first deconvolution layer is connected to the input of the first activation layer, the output of the first activation layer is connected to the input of the convolution layer, and the output of the convolution layer is connected to the input of the second activation layer; the input of the first deconvolution layer and the output of the second activation layer are simultaneously connected to the input of the subtraction layer, the output of the subtraction layer is connected to the input of the second deconvolution layer, and the output of the second deconvolution layer is connected to the input of the third activation layer; the output of the first activation layer and the output of the third activation layer are simultaneously connected to the input of the addition layer, and the output of the addition layer forms the output of the alternating convolution layer; The input of the conversion layer forms the input of the crack segmentation network; the output of the conversion layer is connected to the input of the first self-attention convolutional layer, the output of the first self-attention convolutional layer is connected to the input of the second self-attention convolutional layer, and the output of the second self-attention convolutional layer is connected to the input of the first mixed pooling layer; the output of the first mixed pooling layer is connected to the input of the third self-attention convolutional layer, the output of the third self-attention convolutional layer is connected to the input of the fourth self-attention convolutional layer, and the output of the fourth self-attention convolutional layer is connected to the input of the second mixed pooling layer; the output of the second mixed pooling layer is connected to the input of the fifth self-attention convolutional layer, the output of the fifth self-attention convolutional layer is connected to the input of the sixth self-attention convolutional layer, and the output of the sixth self-attention convolutional layer is connected to the seventh self-attention convolutional layer The output of the seventh self-attention convolution layer is connected to the input of the third mixed pooling layer; the output of the third mixed pooling layer is connected to the input of the eighth self-attention convolution layer, the output of the eighth self-attention convolution layer is connected to the input of the ninth self-attention convolution layer, the output of the ninth self-attention convolution layer is connected to the input of the tenth self-attention convolution layer, and the output of the tenth self-attention convolution layer is connected to the input of the first maximum pooling layer; the output of the first maximum pooling layer is connected to the input of the eleventh self-attention convolution layer, the output of the eleventh self-attention convolution layer is connected to the input of the twelfth self-attention convolution layer, the output of the twelfth self-attention convolution layer is connected to the input of the thirteenth self-attention convolution layer, and the output of the thirteenth self-attention convolution layer is connected to the input of the second maximum pooling layer; The output of the second maximum pooling layer is connected to the input of the first maximum inverse pooling layer; the output of the first maximum inverse pooling layer is connected to the input of the fourteenth self-attention convolution layer, the output of the fourteenth self-attention convolution layer is connected to the input of the fifteenth self-attention convolution layer, the output of the fifteenth self-attention convolution layer is connected to the input of the sixteenth self-attention convolution layer, and the output of the sixteenth self-attention convolution layer is connected to the input of the second maximum inverse pooling layer; the output of the second maximum inverse pooling layer is connected to the input of the seventeenth self-attention convolution layer, the output of the seventeenth self-attention convolution layer is connected to the input of the eighteenth self-attention convolution layer, the output of the eighteenth self-attention convolution layer is connected to the input of the nineteenth self-attention convolution layer, and the output of the nineteenth self-attention convolution layer is connected to the input of the first alternating convolution layer; the first alternating convolution layer The output of the convolution layer is connected to the input of the 20th self-attention convolution layer, the output of the 20th self-attention convolution layer is connected to the input of the 21st self-attention convolution layer, the output of the 21st self-attention convolution layer is connected to the input of the 22nd self-attention convolution layer, and the output of the 22nd self-attention convolution layer is connected to the input of the second alternating convolution layer; the output of the second alternating convolution layer is connected to the input of the 23rd self-attention convolution layer, the output of the 23rd self-attention convolution layer is connected to the input of the 24th self-attention convolution layer, and the output of the 24th self-attention convolution layer is connected to the input of the third alternating convolution layer; the output of the third alternating convolution layer is connected to the input of the 25th self-attention convolution layer, and the output of the 25th self-attention convolution layer is connected to the input of the 26th self-attention convolution layer; The outputs of the 11th, 12th, 13th, 14th, 15th, and 16th self-attention convolutional layers are simultaneously connected to the first scaled attention layer; The outputs of the eighth, ninth, tenth, seventeenth, eighteenth, and nineteenth self-attention convolutional layers are simultaneously connected to the second scaled attention layer; The outputs of the fifth, sixth, seventh, twentieth, twenty-first, and twenty-second self-attention convolutional layers are simultaneously connected to the third scaled attention layer; the outputs of the third, fourth, twenty-third, and twenty-fourth self-attention convolutional layers are simultaneously connected to the fourth scaled attention layer; the outputs of the first, second, twenty-fifth, and twenty-sixth self-attention convolutional layers are simultaneously connected to the fifth scaled attention layer; the outputs of the conversion layer and the outputs of the first to fifth scaled attention layers are simultaneously connected to the input of the splicing layer, the output of the conversion layer is connected to the input of the convolutional layer, and the output of the convolutional layer forms the output of the crack segmentation network; Step 2: Obtain a sample crack data set, and process each sample crack image in the crack data set using the characteristic of sequential changes in pixel intensity to generate an auxiliary image of each sample crack image; Step 3: Use the sample crack image and its auxiliary image to train the crack segmentation network in step 1 to obtain a crack segmentation network model; Step 4: Acquire the crack image to be detected, and process the crack image to be detected using the characteristic of sequential changes in pixel intensity to generate an auxiliary image of the crack image to be detected; Step 5: The crack image to be detected and its auxiliary image are fed into the crack segmentation network model of step 3 to complete crack detection of the crack image to be detected; In the above steps 2 and 4, the specific process of processing the crack image and generating the auxiliary image of the crack image by utilizing the characteristic of sequential change of pixel intensity is as follows: Step 1) For each pixel on the crack image , the crack image and the pixel point exist Distance in direction Pixels are considered as pixels Neighboring pixels ; Step 2) Calculate each pixel of the crack image Its adjacent pixels The relative value of pixel intensity : in, Represents pixel points The pixel intensity, Represents pixel points Neighboring pixels Pixel intensity; Step 3) Exploit The function is applied to each pixel of the crack image Its adjacent pixels The relative value of pixel intensity Normalize and get each pixel of the crack image Its adjacent pixels The normalized pixel intensity relative value ; Step 4) Calculate each pixel of the crack image Its adjacent pixels Position strength : in, is the set mean, is the set variance, is the set intensity weight, Pixel Its adjacent pixels The Euclidean distance from ; Step 5) Calculate each pixel of the crack image Its adjacent pixels The normalized pixel intensity relative value and position strength The mean after multiplication is obtained for each pixel of the crack image Auxiliary pixel intensity ; Step 6) According to each pixel of the crack image The pixel intensity of each pixel is calculated Auxiliary pixel intensity , which constitutes the auxiliary image of the crack image; above , is the number of directions set, , The number of Euclidean distances to be set.
2. The crack detection method based on sequential pixel intensity variation and segmentation network according to claim 1, characterized in that: The mixed pooling layer consists of 1 horizontal pooling layer, 1 vertical pooling layer, 1 average pooling layer and 1 addition layer; The inputs of the horizontal pooling layer, vertical pooling layer, and average pooling layer together form the input of the mixed pooling layer. The inputs of the horizontal pooling layer, vertical pooling layer, and average pooling layer are simultaneously connected to the input of the addition layer, and the output of the addition layer forms the output of the mixed pooling layer.
Citation Information
Patent Citations
Composite degraded image decoupling analysis and restoration method based on cross-branch connection network
CN114266709A
Dam crack intelligent detection method based on unmanned aerial vehicle visual perception and deep learning
CN115880594A