Building structure crack identification method, system and device and storage medium

By introducing a lightweight LWU-Net model into the crack recognition model of building structures, combining the depth separation convolution and simplifying the channel attention mechanism, the problem of large amount of parameters and insufficient recognition accuracy of the existing model is solved, and high-precision crack detection and parameter calculation are realized, which is suitable for lightweight equipment.

CN119942335APending Publication Date: 2025-05-06黄河山
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510030404.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing building structure crack identification model parameters are huge, difficult to deploy on-site, and the recognition accuracy is insufficient in complex backgrounds, especially the extraction capacity of fine cracks is limited.

Method used

A lightweight LWU-Net model is proposed, combining depth separation convolution, space and channel reconstruction convolution, and simplified channel attention mechanism for identification and parameter calculation of building structure cracks. The model enables high-precision crack detection on lightweight devices through preprocessing and mixed loss function optimization.

Benefits of technology

It realizes high-precision identification of cracks that are difficult to detect by mainstream algorithms, overcomes the influence of complex background interference, and can conduct sensitive detection of small cracks on building structures without physical contact, providing technical reserves on lightweight equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942335A_ABST
    Figure CN119942335A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and discloses a building structure crack identification method, system and device and a storage medium, and the method comprises the steps: collecting the image data of a crack disease of a building facade; preprocessing the image data, and dividing the preprocessed image data into a training set and a verification set; constructing a lightweight LWU-Net model, training the model by using the training set, and verifying the performance of the model by using the verification set; the model is used for outputting an identification result of whether the to-be-detected image contains the crack; inputting any to-be-detected image data into the trained model, and outputting an identification result whether the crack is included; and if yes, calculating and outputting actual parameters of the crack according to the image parameters. The method has the advantages that cracks which are not easy to detect by a mainstream algorithm can be detected, and the detection precision is high; and technical reserve is provided for realizing building crack detection on lightweight equipment in the future.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a method, system, device and storage medium for identifying cracks in a building structure. Background Art

[0002] Cracks in building structures are a sign of structural damage or even failure. The causes of cracks are the result of the combined effects of building materials, structural characteristics, and service environment. Structural damage caused by temperature changes and foundation settlement will cause tiny cracks to appear on the surface of the structure, which will then continue to expand and form macro cracks, eventually leading to structural failure. Therefore, timely identification of cracks on the building facade is of great practical engineering significance for understanding the development status of cracks on the building surface and exploring the laws of crack development.

[0003] At present, the traditional detection method is based on manual visual inspection, but its accuracy is affected by the subjectivity of technicians, and the building structures are often relatively tall, so there are certain safety risks during on-site inspection. With the development of computer vision, methods such as edge detection, threshold segmentation and frequency domain detection have been widely used in the field of structural crack detection. However, the detection accuracy of these algorithms will be affected by light intensity, shadows, water stains and complex backgrounds in engineering practice. In recent years, deep learning has demonstrated powerful feature extraction capabilities in the field of disease identification by virtue of its advantages in data analysis. Among them, a series of deep learning algorithms represented by convolutional neural networks have been used for the identification and semantic segmentation of crack diseases, and good recognition results have been obtained, realizing the automatic detection of cracks.

[0004] However, as of now, there are still two technical difficulties in the field of building structure crack identification: first, the mainstream crack identification model has a relatively large number of parameters, and is difficult to deploy on the ground due to the limitations of existing hardware computing power; second, the recognition accuracy of the current model still has room for further improvement, especially for images with complex backgrounds, the existing models are generally insufficient in their ability to extract small cracks.

[0005] Therefore, based on the above technical problems, it is urgently necessary to develop a method, system, device and storage medium for identifying cracks in building structures to solve the above problems. Summary of the invention

[0006] In order to solve the above technical problems, the present invention provides a method, system, device and storage medium for identifying cracks in building structures, which can detect cracks that are difficult to detect with mainstream algorithms and have high detection accuracy; and provides technical reserves for realizing building crack detection on lightweight equipment in the future.

[0007] In a first aspect, the present invention provides a method for identifying cracks in a building structure, comprising the following steps:

[0008] S1, collect image data of cracks on the building facade;

[0009] S2, preprocessing the image data, and dividing the preprocessed image data into a training set and a validation set;

[0010] S3. Construct a lightweight LWU-Net model, train the model using the training set and verify the model performance using the validation set; the model is used to output the recognition result of whether the image to be detected contains cracks; the model is a U-shaped structure formed by several nested layers, and the decoder and encoder of each layer include depthwise separable convolution, spatial and channel reconstruction convolution, and simplified channel attention mechanism;

[0011] S4. Input any image data to be detected into the trained model and output the recognition result of whether it contains cracks; if so, calculate and output the actual parameters of the cracks according to the image parameters.

[0012] Further, in S2, the preprocessing includes:

[0013] S21, removing background occlusion, blur and distortion of image data;

[0014] S22, manually marking the crack position, crack shape and crack size in the remaining image data.

[0015] Furthermore, in the decoder and encoder of each layer, the image data undergoes standard convolution, depthwise separable convolution, standard convolution, simplified channel attention mechanism, spatial and channel reconstruction convolution, simplified channel attention mechanism, standard convolution, depthwise separable convolution and standard convolution in sequence.

[0016] Furthermore, the model parameters are updated using the mixed loss, which is obtained by adding BCE loss, Dice loss, IoU loss and Focal loss.

[0017] Furthermore, the space and channel reconstruction unit includes a space reconstruction module and a channel reconstruction module; the space reconstruction module converts the output data of the previous stage into space refinement features and transmits them to the channel reconstruction unit, and the channel reconstruction unit splits the channel obtained by the space refinement features into a first channel branch and a second channel branch, and the features included in the first channel branch are greater than the features included in the second channel branch;

[0018] Both the first channel branch and the second channel branch are compressed by 1×1 convolution. The spatial refinement features in the first channel branch are subjected to group convolution to obtain a first feature map. The spatial refinement features in the second channel branch are subjected to point-by-point convolution to obtain a second feature map. The feature importance vectors of the first channel branch and the second channel branch are calculated according to the first feature map and the second feature map, and the channel refinement features are calculated according to the feature importance vectors.

[0019] Furthermore, the output feature SCA(X) of the simplified channel attention mechanism is expressed as follows:

[0020] SCA(X)=X·W max-pool (X),

[0021] Among them, max-pool represents the global maximum pooling, W max-pool represents a fully connected layer, and X represents the features of the output of the previous layer.

[0022] Further, in S4, the actual parameters of the crack are calculated and output according to the image parameters, including:

[0023] Get the actual crack size based on the image size;

[0024] The length, width and area of ​​the crack are obtained by extracting the skeleton curve;

[0025] Among them, the crack width is obtained by the watershed method, the crack length is obtained by the segmented superposition method, and the crack area is obtained by the counting method.

[0026] In a second aspect, the present invention provides a system for identifying cracks in a building structure, comprising the following modules:

[0027] Data processing module: collect image data of cracks on the building facade; pre-process the image data and divide the pre-processed image data into a training set and a validation set;

[0028] Computing module: connected to the data processing module, used to build a lightweight LWU-Net model, train the model using the training set and verify the model performance using the validation set; the model is used to output the recognition result of whether the image to be detected contains cracks; the model is a U-shaped structure formed by several nested layers, and the decoder and encoder of each layer include depth-separable convolution, spatial and channel reconstruction convolution, and simplified channel attention mechanism;

[0029] Output module: connected to the calculation module, used to input any image data to be detected into the trained model and output the recognition result of whether it contains cracks; if so, the actual parameters of the cracks are calculated and output according to the image parameters.

[0030] In a third aspect, the present invention provides an electronic device, the electronic device comprising:

[0031] Processor and memory;

[0032] The processor is used to execute the steps of the method of the first aspect by calling the program or instructions stored in the memory.

[0033] In a fourth aspect, the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a program or instruction, wherein the program or instruction enables a computer to execute the steps of the method of the first aspect.

[0034] The present invention has the following technical effects:

[0035] This application constructs the LWU-Net model, which combines deep separable convolution, spatial and channel reconstruction convolution, and simplified channel attention mechanism;

[0036] Depthwise separable convolution can significantly reduce the number of network parameters. When the convolution kernel size is fixed at 3×3, the number of parameters of depthwise separable convolution can be reduced by 9 times compared with traditional convolution. However, the reduction in parameters will reduce the feature extraction capability of the convolution layer. Therefore, spatial and channel reconstruction convolutions are added to each layer of the LWU-Net network to achieve a balance between network lightweighting and feature extraction.

[0037] The spatial and channel reconstruction convolution can reduce the redundancy of the spatial dimension and the channel dimension at the same time. By adopting channel cross reconstruction and channel fusion, the crack extraction performance is guaranteed not to be lost.

[0038] The simplified channel attention mechanism further improves the training efficiency of the network, making the network more focused on feature extraction in the crack area, while ensuring the improvement of training efficiency and suppressing the excessive increase of network parameters;

[0039] This application overcomes the shortcomings of traditional detection methods that are affected by complex background interference and affect recognition accuracy; no physical contact is required and no damage is caused to the concrete; this application is sensitive to small cracks and can detect cracks that are difficult to detect with mainstream algorithms, with high detection accuracy; the introduction of the LWU-Net network provides a technical reserve for the realization of building crack detection on lightweight equipment in the future. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0041] Figure 1 is a flow chart of a method for identifying cracks in a building structure provided by an embodiment of the present invention;

[0042] Figure 2 is a schematic diagram of the structure of a lightweight LWU-Net model provided by an embodiment of the present invention;

[0043] Figure 3 Schematic diagram of the principle of spatial and channel reconstruction convolution provided by an embodiment of the present invention;

[0044] Figure 4 is a schematic diagram of the principle of a simplified channel attention mechanism provided by an embodiment of the present invention;

[0045] Figure 5 is a schematic diagram of crack length calculation provided by an embodiment of the present invention;

[0046] Figure 6 It is a sample diagram of building crack damage provided by an embodiment of the present invention;

[0047] Figure 7 This is a crack disease identification result diagram provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0048] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the scope of protection of the present invention.

[0049] Example 1

[0050] Figure 1 The present invention provides a flow chart of a method for identifying cracks in a building structure.

[0051] See also Figure 1 , including:

[0052] S1, collect image data of cracks on the building facade;

[0053] See also Figure 6 ,UAVs can be used to collect image data and build relevant data sets.,Images of different buildings, positions, angles, and distances should be collected, and the amount of data should be as large as possible to ensure the accuracy of the results.

[0054] S2. Preprocess the image data and divide the preprocessed image data into a training set and a validation set.

[0055] Specifically, preprocessing includes:

[0056] S21, eliminating background-occluded, blurred and distorted image data; screening to eliminate images that are motion blurred, severely distorted or have wall surfaces blocked due to shooting techniques, etc., which may cause large errors in the model results, so as to further improve the accuracy of the model.

[0057] S22, manually marking the crack position, crack shape and crack size in the remaining image data.

[0058] The pre-processed image data is manually annotated, and the starting point, end point, width and shape of the crack are accurately located by formulating unified annotation standards and rules.

[0059] It should be noted that the quality of image annotation will have a significant impact on the accuracy of the model. When annotating, the position, shape and size of the annotated cracks should be consistent with the actual crack area to minimize human bias. In addition, the annotation results can be reviewed again to avoid omissions or missed labels.

[0060] Exemplarily, the preprocessed image data may be used to establish a training set and a validation set in a ratio of 8:2.

[0061] S3. Build a lightweight LWU-Net model, train the model using the training set, and verify the model performance using the validation set; the model is used to output the recognition result of whether the image to be detected contains cracks.

[0062] The lightweight LWU-Net (Lightweight U-Net, LWU-Net) model constructed in this application, that is, the lightweight semantic segmentation network, is an application of supervised learning. Semantic segmentation assigns a semantic label to each pixel in the image, accurately classifies each object area in the image, and specifically predicts the data to be detected by learning the mapping relationship between the input data and the output label. Therefore, manual annotation was performed in the previous S22.

[0063] The model is a U-shaped structure formed by nesting several layers. The decoder and encoder of each layer include depth-wise separable convolution, spatial and channel reconstruction convolution, and simplified channel attention mechanism. For example, see Figure 2 , which can be a U-shaped structure formed by 5 nested layers.

[0064] (1) Depthwise Separable Convolution

[0065] Depth-wise separable convolution can be divided into depth-wise convolution and point-wise convolution. Depth-wise convolution is to perform convolution calculation on each channel of the input layer independently, which saves computational cost, but cannot combine new features generated by various feature maps. Therefore, through point-wise convolution, that is, 1×1 convolution linearly combines the output of depth-wise convolution, the computational cost of depth-wise separable convolution is P SC for:

[0066] P SC =D K ×D K ×C×D F ×DF +C×N×D F ×D F ,

[0067] Among them, C is the dimension of the input image, N is the dimension of the output image, and D K is the size of the convolution kernel, D F is the size of the feature map.

[0068] In one embodiment, for a 3×3 convolution kernel, the computational cost of using depthwise separable convolution is 8 to 9 times lower than that of standard convolution. This application reduces the amount of computation and model size by introducing depthwise separable convolution, solving the problem of high technical cost of complex embedded frameworks.

[0069] (2) Spatial and channel reconstruction convolution

[0070] The spatial and channel reconstruction convolution reduces the network parameters, including the spatial reconstruction unit and the channel reconstruction unit. The spatial reconstruction unit reduces the redundancy of the spatial dimension through separation and reconstruction operations, and the channel reconstruction unit reduces the channel redundancy through the segmentation-conversion-fusion method, thereby improving the model performance.

[0071] The space and channel reconstruction unit includes a space reconstruction module and a channel reconstruction module; the space reconstruction module converts the output data of the previous stage into space refinement features and transmits them to the channel reconstruction unit, and the channel reconstruction unit splits the channel obtained by the space refinement features into a first channel branch and a second channel branch, and the first channel branch contains more features than the second channel branch.

[0072] Both the first channel branch and the second channel branch are compressed by 1×1 convolution. The spatial refinement features in the first channel branch are subjected to group convolution to obtain a first feature map. The spatial refinement features in the second channel branch are subjected to point-by-point convolution to obtain a second feature map. The feature importance vectors of the first channel branch and the second channel branch are calculated according to the first feature map and the second feature map, and the channel refinement features are calculated according to the feature importance vectors.

[0073] See also Figure 3 , the spatial reconstruction unit specifically includes the following steps:

[0074] The output data X of the previous stage is normalized to be mapped to the range of (0, 1), which is specifically achieved by the following formula:

[0075]

[0076] Among them, X out represents the normalized X, μ is the mean, σ is the standard deviation of the input data X, ε is a small constant added for stability calculation, and γ and β are trainable affine transformations.

[0077] X out The relevant weight W γ The expression is:

[0078]

[0079] Among them, γ i Represents the i-th trainable parameter. Then, W is activated by the Sigmoid function. γ Map to the range of (0, 1); set a threshold, and the relevant weights above the threshold are set to 1, as the weight W1 with semantic information; set the relevant weights less than the threshold to 0, as the weight W2 without semantic information. For example, the threshold can be 0.5. The whole process of obtaining weights W1 and W2 can be expressed as W:

[0080] W=Gate(Sigmoid(W γ (GN(X)))).

[0081] Among them, Gate() represents the gate function and GN() represents the sign function. After obtaining W, two weighted feature maps can be obtained: X with weight W1 and rich semantic information. ω11 and X, which has less semantic information and weighted as W2 ω21 ;X ω11 and X ω12 Combined together is X out , X ω21 and X ω22 Combined together, it is also X out In order to reduce spatial redundancy, a cross-reconstruction strategy is adopted to fuse X ω11 and X with fewer features ω22 , and fusion X ω12 and X ω21 Finally, the cross-reconstructed X ω12 With X ω22 The features are connected in series to obtain the spatial refinement feature X ω .

[0082] The channel reconstruction unit includes the following steps:

[0083] Use parameter α (0<α<1) to refine the spatial feature X ω The obtained channel is split into a C channel branch and a (1-α)C channel branch, ie, a first channel branch and a second channel branch.

[0084] In order to reduce the computational cost, 1×1 convolution is used to compress the two channel branches respectively. The spatial refinement feature of the first channel branch is defined as X up, the spatial refinement feature of the second channel branch is defined as X low ; Among them, the spatial refinement feature definition passed by the second channel branch is limited by the coefficient (1-α), and the included features are reduced.

[0085] Compared with the traditional space and channel reconstruction unit, up The transformation stage only uses grouped convolution to extract features instead of the computationally expensive standard convolution to form the combined first feature map Y1, which can be expressed as:

[0086] Y1=M G X up ,

[0087] Among them, M G is the learnable weight matrix of grouped convolution, Y1∈C×h×w, X up ∈αC / r×h×w, h represents the height of X, w represents the width of X, r represents the compression coefficient and r=2.

[0088] For X low In the transformation stage, given that the extracted feature content is small, only 1×1 point-by-point convolution is used to extract hidden information to form a representative second feature map Y2, which can be expressed as:

[0089]

[0090] Among them, M P is the learnable weight matrix of 1×1 point-wise convolution, X low ∈(1-α)C / r×h×w, Y2∈C×h×w.

[0091] Finally, global average pooling is used to collect global spatial information and perform channel statistics S m ∈C×1×1,S m The calculation formula is as follows:

[0092]

[0093] Among them, Pooling represents the pooling operation, Y c (j, k) represents the channel refinement feature of the pixel with height j and width k.

[0094] The channel statistics S1 and S2 are input into the Softmax activation function to calculate the feature importance vectors β1, β2∈C. The final output channel refinement feature Y consists of weighted Y1 and Y2, and the calculation formula is as follows:

[0095] Y=β1Y1+β2Y2.

[0096] (3) Simplify the channel attention mechanism

[0097] The design of the simplified channel attention SCA network architecture is inspired by channel attention (CA). In order to better capture global information and improve computational efficiency, it first compresses spatial information into channels, and then applies multi-layer perception to calculate channel attention and weight feature maps;

[0098] The calculation formula of CA can be expressed as:

[0099] CA(X)=X·sigmoid(W2max(0, W1pool(X))),

[0100] Among them, X represents feature mapping, pool represents global average pooling, sigmoid represents nonlinear activation function, W1 and W2 represent fully connected layers; it can be expressed as:

[0101] CA(X)=X×ψ(X),

[0102] Among them, ψ(X)=sigmoid(W2max(0,W1pool(X))).

[0103] By retaining the global information fusion module and the channel information fusion module of channel attention, a simplified channel attention SCA is established to balance the network lightweight and crack feature extraction capabilities:

[0104] SCA(X)=X·W max-pool (X),

[0105] Among them, X is the feature map, max-pool represents the global maximum pooling, and W max-pool Represents a fully connected layer. The meaning of the formula is explained as follows:

[0106] See also Figure 4 First, the output data X of the previous level is subjected to the maximum pooling operation to achieve feature dimension reduction; secondly, in order to achieve the design goal of lightweight network, the original two layers of 1×1 convolution are reduced to one layer of 1×1 convolution to achieve full connection operation, that is, after the maximum pooling operation, X passes through a layer of 1×1 convolution, the learned distributed feature representation is mapped to the sample label space, and the modified regression unit ReLU and Sigmoid operation are removed; finally, the output data X of the previous level is fused with the convolution X to further enhance the feature extraction capability of the network in the channel dimension.

[0107] In one embodiment, in the decoder and encoder of each layer, the image data undergoes standard convolution, depth-wise separable convolution, standard convolution, spatial and channel reconstruction convolution, standard convolution, depth-wise separable convolution, and standard convolution in sequence.

[0108] In addition, at each layer of the decoder, the feature map of the corresponding layer of the encoder is introduced into the current layer through a skip-layer connection, so that the feature map of the encoder and the feature map of the decoder are spliced, and the spliced ​​feature map is used as the input of the current layer of the decoder for subsequent processing.

[0109] It should be noted that this application does not limit the number of convolutions, which is subject to actual needs.

[0110] This application uses mixed loss to update model parameters, and the mixed loss is obtained by adding BCE loss, Dice loss, IoU loss and Focal loss.

[0111] To further improve the efficiency of network training, this application uses some optimization techniques. For example, all training uses the AdamW optimizer for decoupled weight decay. The AdamW optimizer improves the regularization in Adam by decoupling weight decay from gradient-based updates.

[0112] The formula for exponential decay of weight can be expressed as:

[0113]

[0114] Among them, λ represents the decay rate of the weight at each step, represents the gradient of the tth batch, ρ is the learning rate, θ t represents the training weight at step t.

[0115] In addition, this application introduces a cosine annealing strategy based on "warm restart" to stabilize network training. Before training, several key parameters are set: warmup_factor = 1E-3, warmup_epochs = 1 and end_factor = 1E-6. This strategy increases the learning rate linearly from the default "warmup_factor" value to 1; and uses cosine annealing scheduling to reduce the learning rate to "end_factor". The remaining hyperparameters of training, such as batch size is set to 4, learning rate is set to 1E-5 and cycle rounds are set to 300.

[0116] S4. Input any image data to be detected into the trained model and output the recognition result of whether it contains cracks; if so, calculate and output the actual parameters of the cracks according to the image parameters.

[0117] If cracks are not included, no subsequent calculations are required.

[0118] Extracting geometric parameters of building structure crack images is an important step in evaluating the severity of crack damage and even building structures. However, the size of the image captured by the camera is not consistent with the size in the real world, so the pixel size of the image needs to be calibrated.

[0119] According to the camera imaging principle, the actual target size can be calculated by the following formula:

[0120]

[0121] Among them, the focal length f and the size of the photosensitive element can be obtained by looking up the internal parameters of the camera, and the working distance D is obtained by the image depth information recorded by the Tof camera when the drone uses it to take each crack disease image, that is, the distance between the optical center of the camera carried by the drone and the crack disease.

[0122] After obtaining the size of the photograph, the actual size corresponding to a single pixel can be obtained. When calculating the geometric parameter size of the crack disease, it is only necessary to calculate the number of pixels, and the corresponding actual size can be calculated by the following formula:

[0123] The actual size of the target = the pixels occupied by the target × the size of a single pixel.

[0124] After obtaining the actual crack size according to the image size, the length, width and area of ​​the crack are obtained by extracting the skeleton curve; among them, the crack width is obtained by the watershed method, the crack length is obtained by the segmented superposition method, and the crack area is obtained by the counting method.

[0125] Specifically, the identified cracks are morphologically processed to extract the crack skeleton, and the geometric parameters of the crack part, such as the crack length, width, and area, are calculated. For example, the skeleton curve can be quickly extracted by calling the morphology.skeletonize function of the skimage library in Python.

[0126] In order to more accurately calculate the length, width and area of ​​the bridge crack, it is necessary to extract the crack skeleton curve from the crack area in the image. The crack skeleton curve is composed of points of a single pixel and takes the central axis of the crack target, which can effectively reflect the connectivity and topological structure of the original crack shape. In this process, the center line of the bridge crack target can be obtained by extracting the points of a single pixel, thereby revealing the structural characteristics of the object.

[0127] Calculation of crack width:

[0128] By transforming the edge line of the crack into the local coordinate system of the skeleton line normal vector, and then calculating the tangent line of each point on the skeleton line in the local coordinate system, by constructing the perpendicular line of the tangent line at the point, the perpendicular line intersects the crack edge line at two points (x0, y0) and (x1, y1). Therefore, the calculation formula of the crack width d can be expressed as:

[0129]

[0130] The length of the crack is also calculated based on the skeleton curve of the crack. The calculation of the crack length relies on the method of calculating the distance between pixels in segments and then summing them up. Figure 5 The skeleton curve of the local crack is shown. The black squares represent the crack pixels on the skeleton curve. When the relative positions of the crack pixels on the skeleton are points A and B in the figure, assuming that the coordinates of point A are (x3, y3) and the coordinates of point B are (x2, y2), then their relative distance d is AB It can be expressed as:

[0131]

[0132] The area of ​​structural cracks is calculated using the counting method. The crack image is essentially composed of a large number of pixels. The area where the crack is located is found, and the number of pixels is calculated to determine the area of ​​the crack.

[0133] The processing results of crack damage in this application are compared with those based on the traditional U-Net network and AttU-Net network, see Figure 7 Obviously, the processing result of this application is the best.

[0134] In summary, this application constructs the LWU-Net model, which integrates deep separable convolution, spatial and channel reconstruction convolution, and simplified channel attention mechanism;

[0135] Depthwise separable convolution can significantly reduce the number of network parameters. When the convolution kernel size is fixed at 3×3, the number of parameters of depthwise separable convolution can be reduced by 9 times compared with traditional convolution. However, the reduction in parameters will reduce the feature extraction capability of the convolution layer. Therefore, spatial and channel reconstruction convolutions are added to each layer of the LWU-Net network to achieve a balance between network lightweighting and feature extraction.

[0136] The spatial and channel reconstruction convolution can reduce the redundancy of the spatial dimension and the channel dimension at the same time. By adopting channel cross reconstruction and channel fusion, the crack extraction performance is guaranteed not to be lost.

[0137] The simplified channel attention mechanism further improves the training efficiency of the network, making the network more focused on feature extraction in the crack area, while ensuring the improvement of training efficiency and suppressing the excessive increase of network parameters.

[0138] Example 2

[0139] Based on the above embodiments, the present application provides a system for identifying cracks in building structures, including the following modules:

[0140] Data processing module: collect image data of cracks on the building facade; pre-process the image data and divide the pre-processed image data into a training set and a validation set;

[0141] Computing module: connected to the data processing module, used to build a lightweight LWU-Net model, train the model using the training set and verify the model performance using the validation set; the model is used to output the recognition result of whether the image to be detected contains cracks; the model is a U-shaped structure formed by several nested layers, and the decoder and encoder of each layer include depth-separable convolution, spatial and channel reconstruction convolution, and simplified channel attention mechanism;

[0142] Output module: connected to the calculation module, used to input any image data to be detected into the trained model and output the recognition result of whether it contains cracks; if so, the actual parameters of the cracks are calculated and output according to the image parameters.

[0143] For other details, please refer to Example 1 and will not be described in detail here.

[0144] Example 3

[0145] Based on the aforementioned embodiments, the present application provides an electronic device, which includes: a processor and a memory; the processor is used to execute the steps of the method in Embodiment 1 by calling a program or instruction stored in the memory.

[0146] Example 4

[0147] A computer-readable storage medium stores a program or instruction, wherein the program or instruction enables a computer to execute the steps of the method in Example 1.

[0148] It should be noted that the terms used in the present invention are only for describing specific embodiments, rather than limiting the scope of the present application. As shown in the present specification, unless the context clearly indicates an exception, the words "one", "a", "a kind of" and / or "the" do not specifically refer to the singular, but may also include the plural. The terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method or device. In the absence of more restrictions, the elements defined by the sentence "include one..." do not exclude the presence of other identical elements in the process, method or device including the elements.

[0149] It should also be noted that the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. Unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", etc. should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be an indirect connection through an intermediate medium, or it can be a connection between the two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0150] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.

Claims

1. A method for identifying cracks in a building structure, characterized in that: The steps include: S1, collect image data of cracks on the building facade; S2, preprocessing the image data, and dividing the preprocessed image data into a training set and a verification set; S3, constructing a lightweight LWU-Net model, training the model using the training set and verifying the model performance using the verification set; the model is used to output a recognition result of whether the image to be detected contains cracks; wherein the model is a U-shaped structure formed by nesting several layers, and the decoder and encoder of each layer include depthwise separable convolution, spatial and channel reconstruction convolution, and simplified channel attention mechanism; S4. Input any image data to be detected into the trained model and output a recognition result of whether it contains cracks; if so, calculate and output the actual parameters of the cracks according to the image parameters.

2. A method for identifying cracks in a building structure according to claim 1, characterized in that: In S2, the preprocessing includes: S21, removing background occlusion, blur and distortion of image data; S22, manually marking the crack position, crack shape and crack size in the remaining image data.

3. A method for identifying cracks in a building structure according to claim 1, characterized in that: In the decoder and encoder of each layer, the image data undergoes standard convolution, depthwise separable convolution, standard convolution, simplified channel attention mechanism, spatial and channel reconstruction convolution, simplified channel attention mechanism, standard convolution, depthwise separable convolution and standard convolution in sequence.

4. A method for identifying cracks in a building structure according to claim 1, characterized in that: The model parameters are updated using a mixed loss, which is obtained by adding BCE loss, Dice loss, IoU loss and Focal loss.

5. A method for identifying cracks in a building structure according to claim 1, characterized in that: The space and channel reconstruction unit includes a space reconstruction module and a channel reconstruction module; the space reconstruction module converts the output data of the previous stage into space refinement features and transmits them to the channel reconstruction unit; the channel reconstruction unit splits the channel obtained by the space refinement features into a first channel branch and a second channel branch, and the features contained in the first channel branch are greater than the features contained in the second channel branch; The first channel branch and the second channel branch are both compressed by 1×1 convolution, the spatial refinement features in the first channel branch are subjected to group convolution to obtain a first feature map, the spatial refinement features in the second channel branch are subjected to point-by-point convolution to obtain a second feature map, the feature importance vectors of the first channel branch and the second channel branch are calculated according to the first feature map and the second feature map, and the channel refinement features are calculated according to the feature importance vectors.

6. A method for identifying cracks in a building structure according to claim 1, characterized in that: The output feature SCA(X) of the simplified channel attention mechanism is expressed by the following formula: SCA(X)=X·W max-pool (X), Among them, max-pool represents the global maximum pooling, W max-pool represents a fully connected layer, and X represents the features of the output of the previous layer.

7. A method for identifying cracks in a building structure according to claim 1, characterized in that: In S4, the actual parameters of the crack are calculated and output according to the image parameters, including: Get the actual crack size based on the image size; The length, width and area of ​​the crack are obtained by extracting the skeleton curve; Among them, the crack width is obtained by the watershed method, the crack length is obtained by the segmented superposition method, and the crack area is obtained by the counting method.

8. A system for identifying cracks in a building structure, used to execute the method for identifying cracks in a building structure as claimed in any one of claims 1 to 7, characterized in that: Includes the following modules: Data processing module: collects image data of cracks on the facade of a building; pre-processes the image data, and divides the pre-processed image data into a training set and a verification set; Computing module: connected to the data processing module, used to construct a lightweight LWU-Net model, train the model using the training set and verify the model performance using the verification set; the model is used to output a recognition result of whether the image to be detected contains cracks; wherein the model is a U-shaped structure formed by nesting several layers, and the decoder and encoder of each layer include depth-separable convolution, spatial and channel reconstruction convolution and simplified channel attention mechanism; Output module: connected to the calculation module, used to input any image data to be detected into the trained model and output the recognition result of whether it contains cracks; if so, calculate and output the actual parameters of the cracks according to the image parameters.

9. An electronic device, characterized in that: The electronic device comprises: Processor and memory; The processor is used to execute the steps of the method for identifying cracks in a building structure as described in any one of claims 1 to 7 by calling the program or instruction stored in the memory.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program or instruction, and the program or instruction enables a computer to execute the steps of the method for identifying cracks in a building structure as claimed in any one of claims 1 to 7.