A Wood Crack Detection Method Based on Semantic Segmentation Network

By adopting a semantic segmentation network-based method in wood crack detection, combining position attention mechanism, feature enhancement mechanism and residual module, the problems of low detection accuracy and low efficiency in the prior art are solved, and automated detection and robustness are achieved.

CN115908325BActive Publication Date: 2025-06-03FUZHOU UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211461477.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-17
Publication Date
2025-06-03
Estimated Expiration
2042-11-17

AI Technical Summary

Technical Problem

The prior art has problems of low accuracy, low efficiency and insensitive to detection of fine cracks in wood cracks, especially in the case of unstable wood image brightness and complex background texture.

Method used

The detection method based on semantic segmentation network is adopted, U-Net is used as the basic network structure, and a position attention mechanism, feature enhancement mechanism and residual module with hollow convolution are designed to highlight the crack position information, enhance the detailed information of small cracks and extract information of long cracks and wide cracks.

Benefits of technology

Automatic detection of wood cracks is realized, labor costs are reduced, detection robustness and recognition rate are improved, and detection rate is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115908325B_ABST
    Figure CN115908325B_ABST
Patent Text Reader

Abstract

The present invention relates to a wood crack detection method based on a semantic segmentation network. First, a position attention mechanism is proposed to highlight the position of wood cracks; at the same time, a feature enhancement mechanism is designed to enhance more detailed information of small wood cracks. In addition, residual blocks are used to fuse multi-scale receptive fields to obtain more crack area information through larger receptive fields. The beneficial effect of the present invention is to reduce the influence of other interference factors, be applicable to crack detection of different scales, and improve the accuracy of wood crack detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wood defect detection, and particularly to a wood crack detection method based on a semantic segmentation network. Background Art

[0002] Wood board cracks are typical defects that affect the quality of wood products. Therefore, during the production process of wooden furniture, it is necessary to detect and cut wood cracks. Generally, wood cracks are marked by experienced workers with a fluorescent pen and then used for wood cutting decisions. Manual processing is inefficient, and human subjectivity often leads to low accuracy.

[0003] In the wood industry, crack detection is one of the most challenging tasks in wood defect detection. During the industrial production process, wood boards often contain stains, and due to unstable lighting, the captured wood images have different brightnesses, reducing the accuracy of wood crack detection. On the other hand, the fine cracks on the wood and some cracks similar to normal areas make the segmentation of wood cracks more difficult.

[0004] Using machine learning methods to detect cracks on the surface of wood raw materials can effectively avoid the influence of human subjective factors, and at the same time can effectively improve the accuracy of wood defect detection and improve the detection efficiency. Therefore, developing an automatic wood crack detection method is of great significance for the wood industry.

[0005] The Chinese patent application number is: CN201910471773.0, and the name is: Method for detecting surface defects of solar cells. This method performs normalization processing on the picture to obtain a binary image of the solar cell surface image, establishes a deep learning model based on the convolutional neural network structure to form a deep belief network. The deep belief network is trained through the deep learning model, and the trained deep belief network is used to detect the surface defects of solar cells on the binary image of the test set, and the deep belief network outputs the detection result. This method only performs a preliminary binary classification on the solar surface defects and cannot obtain more defect size information and defect location information.

[0006] The Chinese patent application number is: CN201910607197.8, and the name is: A visual defect detection method based on a deep convolutional neural network. This method is used for detecting surface defects of photovoltaic cells. By creatively integrating the structure decoupling function of SEF into CNN while retaining the feature extraction ability of the common convolutional layer, it strengthens the effectiveness and accuracy of the model in extracting complex crack defect features, realizes the decoupling of features and the background, and can effectively solve the problems of complex background textures, diverse crack defect features, and random shapes on the surface of the battery cells. This method is not sensitive enough to tiny defects, and when there are fine cracks, it is easy to miss detections. Summary of the Invention

[0007] The object of the present invention is to provide a wood crack detection method based on a semantic segmentation network. Taking U-Net as the basic network structure, a position attention mechanism and a feature enhancement mechanism are designed, and a residual module with dilated convolution is introduced, which can highlight the crack position information, enhance the detail information of small cracks, and extract the information of long cracks and wide cracks. When applying the crack detection method of the present invention to the detection of wood surface cracks, the extraction of wood cracks can be effectively realized.

[0008] To achieve the above object, the technical solution of the present invention is: a wood crack detection method based on a semantic segmentation network, and the specific steps are as follows:

[0009] The first step, image acquisition and preprocessing:

[0010] 1-1. Image acquisition: Obtain wood images through a line array camera;

[0011] 1-2. Image cutting: Cut the wood image into equal-width small squares according to the width of the wood image;

[0012] 1-3. Sample set production: Manually add pixel-level labels to the wood images with cracks to make a wood crack sample set;

[0013] The second step, building the WCU-Net structure model:

[0014] Build the WCU-Net structure based on the U-Net network structure. The model includes the basic network U-Net, a position attention module that highlights the wood crack position, a feature enhancement module that enhances small crack information, and a residual block that can obtain and fuse multi-scale receptive fields. The specific structure is as follows:

[0015] 2-1. Basic network U-Net:

[0016] Basic network encoding part: For the encoder part, two convolutional layers (kernel size 3×3) are used to obtain features at different scales, and a max pooling layer with a stride of 2 is used for downsampling. After each convolutional layer, batch normalization (BN) and rectified linear unit activation function (ReLU) operations are performed;

[0017] Basic network decoding part: In the decoding process, an Invertedblock composed of depth convolution and pointwise convolution is used. The first pointwise convolution halves the number of channels, and the remaining convolution operations keep the number of channels unchanged. For feature upsampling, the image size is gradually restored through transposed convolution. In the last layer, Sigmoid is selected as the activation function;

[0018] 2-2. Position attention mechanism:

[0019] 1) Let the feature X∈RC×H×W Input two convolutional layers with kernel sizes of 1×W and H×1 respectively, and generate A∈R H×1 and B∈R 1×W two vectors;

[0020] 2) Perform a matrix multiplication operation between A and B, and use a softmax layer to obtain the position attention map S∈R H×W ;

[0021] 3) Multiply the original feature X by the position attention map S, and then perform element-wise summation with the feature X to generate the final enhanced feature map Y∈R C×H×W 。

[0022] 2-3. Feature enhancement mechanism:

[0023] (1) Utilize the low-level semantic feature map and the high-level semantic feature map to generate two position attention maps α and β according to the position attention mechanism in step 2-2.

[0024] (2) After performing the operation of (1-α) on the position attention map α, perform bilinear interpolation upsampling on the result;

[0025] (3) Multiply the result obtained in step (2) with the corresponding elements of the position attention map β and the low-level semantic feature map;

[0026] (4) Add the result obtained in step (3) to the corresponding elements of the low-level semantic feature map to obtain the final enhanced output;

[0027] 2-4. Residual structure: Introduce residual blocks to obtain a larger receptive field. The used residual blocks are composed of connecting 3 residual units, and a convolution with a kernel size of 3×3 is selected, and the dilation rate is set to 4;

[0028] The third step, training of the wood crack detection model:

[0029] 3-1. Parameter initialization settings: During training, the image size is adjusted to Hi×Wi. The number of training times, the initial learning rate, and the batch size are also set to a, b, and c respectively. The cosine annealing learning rate adjustment method is used to adjust the learning rate, and the Adam optimizer is selected. During the training phase, perform image enhancement operations such as random cropping, random brightness, and random rotation randomly;

[0030] 3-2. Start training: Input the training sample set into the WCU-Net structure network;

[0031] 3-3. Update parameters: Calculate the Dice loss with the corresponding annotation map and perform backpropagation training to update the weight parameters in the WCU-Net network to reduce the loss value. The specific loss function formula is as follows:

[0032]

[0033] where p n and g n represent the predicted value and the ground truth respectively, and the minimum value ε is used to avoid a zero denominator;

[0034] 3-4. Output model: Repeat steps 3-2 and 3-3. When the number of training times reaches the set value, stop training and output the model with the minimum loss value.

[0035] Fourth step, use the trained model for wood surface crack detection.

[0036] Compared with the prior art, the present invention has the following beneficial effects:

[0037] (1) The present invention proposes a wood crack detection method based on a semantic segmentation network, realizing automatic detection of wood cracks and greatly reducing the labor cost.

[0038] (2) A position attention mechanism for highlighting the position of wood cracks is designed, which can tolerate the interference of other factors on crack detection and improve the robustness of wood crack detection.

[0039] (3) A feature enhancement mechanism is designed, which effectively improves the recognition rate of crack detection by selectively obtaining and enhancing more information on small cracks.

[0040] (4) The residual block is used to fuse multi-scale receptive fields to obtain more crack region information and improve the detection rate of large wood crack regions. Description of the Drawings

[0041] Figure 1 is the network structure diagram of the WCU-Net described in the present invention;

[0042] Figure 2 is the structure diagram of the Invertedblock described in the present invention;

[0043] Figure 3 is the generation mode diagram of the position attention map described in the present invention;

[0044] Figure 4 is the structure diagram of the position attention mechanism described in the present invention;

[0045] Figure 5 is the structure diagram of the feature enhancement mechanism described in the present invention;

[0046] Figure 6 is the structure diagram of the residual module described in the present invention. Detailed Embodiment

[0047] The technical solution of the present invention will be specifically described below with reference to the accompanying drawings.

[0048] A method for detecting wood cracks based on a semantic segmentation network according to the present invention comprises the following specific steps:

[0049] The first step, image acquisition and preprocessing:

[0050] 1-1. Image acquisition: Obtain wood images through a line array camera;

[0051] 1-2. Image cutting: Cut the wood image into small squares of equal width according to the width of the wood image;

[0052] 1-3. Sample set production: Manually add pixel-level labels to the wood images with cracks to produce a wood crack sample set. A total of 729 images are obtained from step 1-1, 549 of which are randomly selected for network training and 180 images for testing;

[0053] The second step, building the WCU-Net structure model:

[0054] Build the WCU-Net structure based on the U-Net network structure, as Figure 1 shown. The model includes a basic network U-Net, a position attention module that highlights the position of wood cracks, a feature enhancement module that enhances small crack information, and a residual block that can obtain and fuse multi-scale receptive fields. The specific structure is as follows:

[0055] 2-1. Basic network U-Net:

[0056] Basic network encoding part: For the encoder part, two convolutional layers (kernel size 3×3) are used to obtain features at different scales, and a max pooling layer with a stride of 2 is used for downsampling. After each convolutional layer, batch normalization (BN) and rectified linear unit activation function (ReLU) operations are performed;

[0057] Basic network decoding part: In the decoding process, an Invertedblock composed of depthwise convolution and pointwise convolution is used, as Figure 2 shown. The pointwise convolution in the first layer halves the number of channels, and the remaining convolution operations keep the number of channels unchanged. For feature upsampling, the size of the image is gradually restored through transposed convolution. In the last layer, Sigmoid is selected as the activation function;

[0058] 2-2. Position attention mechanism:

[0059] 1) As Figure 3 , input the feature X∈R C×H×W into two convolutional layers with a kernel size of 1×W and H×1 respectively to generate A∈R H×1and B ∈ R 1×W Two vectors

[0060] 2) Perform a matrix multiplication operation between A and B, and use a softmax layer to obtain the position attention map S ∈ R H×W ;

[0061] 3) Multiply the original feature X by the position attention map S, and then perform element-wise summation with the feature X to generate the final enhanced feature map Y ∈ R C×H×W ; As Figure 4 shown

[0062] 2-3. Feature enhancement mechanism:

[0063] (1) As Figure 5 shown, use the low-level semantic feature map and the high-level semantic feature map to generate two position attention maps α and β according to the position attention mechanism in step 2-2

[0064] (2) After performing the operation of (1-α) on the position attention map α, perform bilinear interpolation upsampling on the result

[0065] (3) Multiply the result obtained in step (2) with the corresponding elements of the position attention map β and the low-level semantic feature map

[0066] (4) Add the result obtained in step (3) to the corresponding elements of the low-level semantic feature map to obtain the final enhanced output

[0067] 2-4. Residual structure: Introduce a residual block to obtain a larger receptive field. As Figure 6 shown, the used residual block is composed of connecting 3 residual units, and a convolution with a convolution kernel size of 3×3 is selected, and the dilation rate is set to 4

[0068] The third step, training of the wood crack detection model:

[0069] 3-1. Parameter initialization settings: During training, the image size is adjusted to 416×416. The number of training times, the initial learning rate, and the batch size are also set to 300, 5e-4, and 4 respectively. The cosine annealing learning rate adjustment method is used to adjust the learning rate, and the Adam optimizer with a weight decay of 5e-4 is used. During the training phase, image enhancement operations such as random cropping, random brightness, and random rotation are performed randomly

[0070] 3-2. Start training: Input the training sample set into the WCU-Net structure network

[0071] 3-3. Update parameters: Calculate the Dice loss with the corresponding annotation map and perform backpropagation training to update the weight parameters in the WCU-Net network, reducing the loss value. The specific loss function formula is as follows:

[0072]

[0073] where p n and g n represent the predicted value and the ground truth respectively, and the minimum value ε is used to avoid a zero denominator;

[0074] 3-4. Output the model: Repeat steps 3-2 and 3-3. When the number of training times reaches the set value, stop training and output the model with the minimum loss value;

[0075] Fourth step, use the trained model for wood surface crack detection.

[0076] In this example, the above wood crack defect detection model and other object detection method models are respectively trained and tested using the training set and the test set, and several metrics including precision (Pr), recall (Re), F1-score (F1), and IoU value are introduced for model evaluation. The experimental results and comparisons are shown in Table 1:

[0077] Table 1

[0078]

[0079]

[0080] Analyzing the above experimental results, it can be seen that: compared with detection methods such as U-Net, PSPNet, and DeepLabV3+, the detection model using the WCU-Net network structure has significantly improved Re value, F1 value, and IoU value, except that the Pr value is slightly lower than that of U-Net, further demonstrating the effectiveness of the present invention.

[0081] The above are the preferred embodiments of the present invention. All changes made according to the technical solution of the present invention that do not exceed the scope of the technical solution of the present invention in terms of the functions and effects produced belong to the protection scope of the present invention.

Claims

1. A wood crack detection method based on a semantic segmentation network, characterized in that, it includes the following steps: The first step, image acquisition and preprocessing; The second step, building the WCU-Net structure model; The third step, training the wood crack detection model; The fourth step, using the trained model for wood surface crack detection; The specific implementation of the first step is as follows: 1-1. Image acquisition: Obtain wood images through a line array camera; 1-2. Image cutting: According to the width of the wood image, crop the wood image into small equal-width squares; 1-3. Sample set production: Manually add pixel-level labels to the wood images with cracks to make a wood crack sample set; In the second step, build the WCU-Net structure model based on the U-Net network structure; The WCU-Net structure model includes the basic network U-Net, a position attention module that highlights the position of wood cracks, a feature enhancement module that enhances small crack information, and a residual block that can obtain and fuse multi-scale receptive fields, specifically as follows: 2-1. Basic network U-Net: Basic network encoding part: Use two convolutional layers to obtain features at different scales, with a kernel size of 3×3, and use a max pooling layer with a stride of 2 for downsampling; After each convolutional layer, perform batch normalization BN and the rectified linear unit activation function ReLU operations; Basic network decoding part: Use Invertedblock composed of depth convolution and pointwise convolution; The pointwise convolution in the first layer halves the number of channels, and the remaining convolution operations keep the number of channels unchanged; For feature upsampling, gradually restore the image size through transposed convolution; In the last layer, select Sigmoid as the activation function; 2-2. Position attention mechanism: 1) Input the feature \(X\in\mathbb{R}\) C×H×W into two convolutional layers with kernel sizes of \(1\times W\) and \(H\times1\) respectively, generating two vectors \(A\in\mathbb{R}\) H×1 and \(B\in\mathbb{R}\) 1×W respectively; 2) Perform a matrix multiplication operation between A and B, and use a softmax layer to obtain the positional attention map S ∈ R H×W ; 3) Multiply the original feature X by the position attention map S, and then perform element-wise summation with the feature X to generate the final enhanced feature map Y ∈ R C×H×W ; 2-3. Feature enhancement mechanism: (1) Utilize the low-level semantic feature map and the high-level semantic feature map to generate two position attention maps α and β according to the position attention mechanism in 2-2; (2) After performing the operation of (1-α) on the position attention map α, then perform bilinear interpolation upsampling on its result; (3) Multiply the result obtained in step (2) with the corresponding elements of the position attention map β and the low-level semantic feature map; (4) Add the corresponding elements of the result obtained in step (3) to the low-level semantic feature map to obtain the final enhanced output; 2-4. Residual structure: Introduce a residual block to obtain a larger receptive field; The used residual block is composed of connecting 3 residual units, and a convolution with a kernel size of 3×3 is selected, and the dilation rate is set to 4; The specific implementation of the third step is as follows: 3-1. Parameter initialization settings: During training, the image size is adjusted to Hi×Wi; The number of training times, the initial learning rate, and the batch size are also set to a, b, and c respectively; Use the cosine annealing learning rate adjustment method to adjust the learning rate, and select the Adam optimizer; During the training phase, perform random cropping, random brightness, and random rotation image enhancement operations randomly; 3-2. Start training: Input the training sample set into the WCU-Net structure model; 3-3. Update parameters: Calculate the Dice loss with the corresponding annotation map and perform backpropagation training to update the weight parameters in the WCU-Net structure model, so as to reduce the loss value. The specific loss function formula is as follows: where p n and g n represent the predicted value and the true value respectively, and the minimum value ε is used to avoid a zero denominator; 3-4. Output the model: Repeat steps 3-2 and 3-3. When the number of training times reaches the set value, stop training and output the model with the minimum loss value.

Citation Information

Patent Citations

  • Solar cell surface defect detection method

    CN110349120A

  • A visual defect detection method using deep convolutional neural networks

    CN110610475B

  • Building facade crack detection method and system based on unmanned aerial vehicle

    CN114705689A