Wood Component Crack Detection Model, Method and System under Complex Background Based on Improved YOLOv11n
By improving the neck network of the YOLOv11n model and introducing attention mechanism, the problem of low crack detection accuracy of wood components under complex background is solved, and high-precision crack detection and recognition is achieved.
Patent Information
- Application Number
- CN202510353330.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-25
AI Technical Summary
The prior art is difficult to detect cracks in wooden components with high precision in complex backgrounds, especially small-scale cracks, with low recognition accuracy.
A crack detection model for wood component under complex background based on improved YOLOv11n is proposed. By improving the neck network, introducing residual branching and attention mechanism, the feature expression ability is enhanced, and the PIoU2 loss function is used to improve the detection accuracy.
It improves the crack detection accuracy of wooden components in complex backgrounds, enhances the crack representation ability, reduces information loss, and achieves efficient detection and identification.
Smart Images

Figure CN119863468B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of civil engineering prefabricated structure detection and computer technology, and in particular to a wooden member crack detection model, method and system based on improved YOLOv11n in a complex background. Background Art
[0002] Cracks are widely distributed in infrastructure such as buildings, roads and bridges, which not only affect the appearance but also threaten the stability and safety of the structure. Therefore, timely and effective crack detection is crucial for ensuring the safety of building structures and extending their service life. Traditional building structure crack detection techniques include manual detection and machine detection. Manual detection methods are inefficient and inaccurate, while machine detection methods such as infrared and thermal imaging, ultrasonic detection, laser detection and ray detection are more accurate, but the equipment cost is high and the operation process is complex.
[0003] Components of concrete structures, roads and bridges are relatively large in size, and it is easy to obtain photos that contain the entire crack characteristics and have a single environmental background. For wood, especially wooden structural members, they are not only relatively small in size, but also the crack shape is relatively close to the wood grain characteristics. When the photo contains the entire crack characteristics, the environmental background is more complex. Existing research on wood crack identification mostly focuses on photos of small-scale cracks, and there is less research on complex background situations.
[0004] With the development of artificial intelligence technology, image recognition technology has become mature and began to be used for object detection in engineering. However, existing artificial intelligence detection technologies are difficult to detect small-scale cracks in complex backgrounds, and there is room for further optimization in the detection of wooden member cracks. Summary of the Invention
[0005] In order to overcome the defect that it is difficult to detect wooden member cracks in the above-mentioned prior art, the present invention proposes a training method for a wooden member crack detection model based on improved YOLOv11n in a complex background, which pays attention to the context dependence of image features and multi-dimensional features, so as to suppress the interference of complex backgrounds, enhance crack representation and improve the accuracy of wooden member crack detection.
[0006] A training method for a crack detection model of wooden components in complex backgrounds based on improved YOLOv11n. First, the Neck network of YOLOv11n is replaced. The new Neck network includes a first convolutional network, a second convolutional network, a third convolutional network, a fourth convolutional network, and a first sampling network, a first splicing module, a second sampling network, a second splicing module, a third splicing module, a fifth convolutional network, a fourth splicing module, a sixth convolutional network, and a fifth splicing module connected in sequence. The inputs of the first, second, third, and fourth convolutional networks are respectively connected to the outputs of four network layers from shallow to deep of the Backbone network. The outputs of the first, second, third, and fourth convolutional networks are respectively connected to the inputs of the third, second, fourth, and fifth splicing modules. The output of the second convolutional network is also connected to the input of the third splicing module, the output of the third convolutional network is also connected to the input of the first splicing module, and the output of the fourth convolutional module is also connected to the input of the first sampling network. The output of the first splicing module is also connected to the input of the fourth splicing module.
[0007] Then, the improved model is made to perform machine learning on a dataset of wooden component images annotated with cracks.
[0008] Preferably, the first splicing module, the second splicing module, the third splicing module, the fourth splicing module, and the fifth splicing module have the same structure, and are all composed of a weighted splicing unit BiCCT and a C3k2 module connected front and back. BiCCT includes a Triplet Attention layer, a Concat-T layer, a first dimension stacking layer, and an activation layer connected in sequence. The first dimension stacking layer stacks the output of the Concat-T layer and the input of BiCCT in dimensions.
[0009] Preferably, the improvement to YOLOv11n also includes: replacing the C3k2 module in the splicing module with a C3k2-Faster-CAA module. The C3k2-Faster-CAA module includes an input convolutional layer, a Split function layer, one or more Faster-CAA modules, a second dimension stacking layer, and an output convolutional layer connected in sequence. The second dimension stacking layer is used to stack the output of the Split function layer and the outputs of each Faster-CAA module in dimensions. The Faster-CAA module introduces an attention mechanism and a residual branch to capture the context dependence between distant pixels.
[0010] Preferably, the Faster-CAA module includes a partial convolutional layer, a first pointwise convolutional layer, a second pointwise convolutional layer, an attention layer, and a third dimension stacking layer connected in sequence. The input of the third dimension stacking layer is also connected to the input of the Faster-CAA module.
[0011] Preferably, the C3k2 module in the Backbone network is replaced with a C3k2-Faster-CAA module.
[0012] Preferably, the input of the first convolutional network is connected to the third layer of the backbone network, the input of the second convolutional network is connected to the fifth layer of the backbone network, the input of the third convolutional network is connected to the seventh layer of the backbone network, and the input of the fourth convolutional network is connected to the eleventh layer of the backbone network.
[0013] Preferably, the loss function adopted during the model training process is the PIoU2 loss.
[0014] A method for detecting cracks in wooden components proposed by the present invention uses the training method of the crack detection model for wooden components in complex backgrounds based on the improved YOLOv11n to obtain a crack detection model; the wooden component to be detected is photographed, and the photo is input into the crack detection model.
[0015] Preferably, the photo of the wooden component to be detected is input into the crack detection model for processing after image enhancement.
[0016] A system for detecting cracks in wooden components proposed by the present invention includes a memory and a processor. A computer program is stored in the memory, and the processor is connected to the memory. The processor is used to execute the computer program to implement the method for detecting cracks in wooden components in complex backgrounds based on the improved YOLOv11n.
[0017] The advantages of the present invention are as follows:
[0018] (1) The training method of the crack detection model for wooden components in complex backgrounds based on the improved YOLOv11n proposed by the present invention improves the neck network of YOLOv11n, introduces a residual branch on the original BiFPN structure, and combines convolutional operations to enhance the feature expression ability, avoid information loss, improve the image feature judgment ability in complex backgrounds, and thus improve the accuracy of wooden structure crack detection.
[0019] (2) The weighted splicing unit BiCCT proposed by the present invention adjusts the weights of multiple input feature maps, adjusts the weights of each channel of the feature maps, then recombines the feature maps by combining the product of the weights of each feature map and the channel weights, improves the feature expression of the feature maps, realizes feature map enhancement, and superimposes the enhanced feature maps with the original input to realize the context echo of the feature maps, thereby flexibly enhancing the representation of the crack area and suppressing the interference of complex backgrounds at the same time.
[0020] (3) The present invention introduces a triple attention mechanism through Triplet Attention to capture and interact with the features of the channel height, channel width, and spatial dimension, and then combines the activation function to enhance the non-linear expression ability, enhance the adaptability of the network to complex crack morphologies, thereby minimizing the loss of crack information and effectively balancing the computational efficiency and accuracy.
[0021] (4) The C3k2-Faster-CAA module of the present invention enhances the feature modeling channel and spatial interdependence through the nesting and residual combination of multiple Faster-CAA modules, thereby improving the processing ability of crack targets in complex backgrounds and enhancing the detection accuracy.
[0022] (5) The Faster-CAA module of the present invention strengthens the features through the cooperation of partial convolution and pointwise convolution, reduces the amount of calculation and memory access, alleviates the computing burden on edge devices, and greatly improves the detection efficiency. The Faster-CAA module introduces an attention mechanism and a residual branch to capture the context dependence between distant pixels.
[0023] (6) The method for detecting cracks in wooden components in complex backgrounds proposed by the present invention, which uses the crack detection model proposed by the present invention for identifying cracks in wooden components, is fast and accurate, and solves the problem that it is difficult to identify cracks in wooden components with high precision due to the complex background in architecture. The present invention has excellent comprehensive performance and low hardware requirements, and becomes an ideal solution suitable for crack detection in wooden structures in complex backgrounds. The improved model proposed by the present invention not only performs outstandingly in terms of detection accuracy, but also shows excellent performance in terms of computing efficiency and hardware adaptability, providing strong technical support for practical applications. Description of the Drawings
[0024] Figure 1 is the structure of the existing YOLOv11n;
[0025] Figure 2 is the structure of YOLOv11n-B given by the present invention;
[0026] Figure 3 is the BiFPN-T structure given by the present invention;
[0027] Figure 4 is the BiCCT structure given by the present invention;
[0028] Figure 5 is the structure of YOLOv11n-BC given by the present invention;
[0029] Figure 6 is the structure of C3k2-Faster-CAA given by the present invention;
[0030] Figure 7 is the structure of Faster-CAA given by the present invention;
[0031] Figure 8 is the comparison diagram of the detection effects before and after improvement in the embodiment;
[0032] Figure 9 It is the flowchart of the model training method in the embodiment;
[0033] Figure 10 It is the P-R curve of the ablation experiment in the embodiment;
[0034] Figure 11 It is the comparison of the model evaluation indexes in the embodiment;
[0035] Figure 12 It is the comparison of the detection category indexes in the embodiment. Detailed implementation manners
[0036] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0037] A complex-background wooden component crack detection model based on improved YOLOv11n proposed by the present invention, hereinafter referred to as the crack detection model for short. The structure of this crack detection model is improved on the basis of YOLOv11n.
[0038] The present invention proposes multiple improvement methods for YOLOv11n, so as to obtain a crack detection model for accurately detecting cracks in wooden components. The traditional YOLOv11n model is as Figure 1 shown.
[0039] Referring to Figure 2 , the first improved model is denoted as YOLOv11n-B. Compared with YOLOv11n, the neck network of the model YOLOv11n-B is improved.
[0040] The neck network of the model YOLOv11n-B includes a first convolutional network, a second convolutional network, a third convolutional network, a fourth convolutional network, and a first sampling network, a first splicing module, a second sampling network, a second splicing module, a third splicing module, a fifth convolutional network, a fourth splicing module, a sixth convolutional network, and a fifth splicing module connected in sequence; the outputs of the third splicing module, the fourth splicing module, and the fifth splicing module are respectively connected to 3 splitting heads of the head network (Head).
[0041] The first convolutional network, the second convolutional network, the third convolutional network, and the fourth convolutional network are used as the input of the neck network connecting the backbone network, and the inputs of the first convolutional network, the second convolutional network, the third convolutional network, and the fourth convolutional network are respectively connected to the outputs of four network layers of the backbone network from shallow to deep.
[0042] Specifically, the input of the first convolutional network is connected to the output of the 3rd layer C3k2 of the backbone network, the input of the second convolutional network is connected to the output of the 5th layer C3k2 of the backbone network, the input of the third convolutional network is connected to the output of the 7th layer C3k2 of the backbone network, and the input of the fourth convolutional network is connected to the output of the 11th layer C3PSA of the backbone network.
[0043] The output of the first convolutional network is connected to the input of the third splicing module; the output of the second convolutional network is respectively connected to the input of the second splicing module and the input of the third splicing module, the output of the third convolutional network is respectively connected to the input of the first splicing module and the input of the fourth splicing module; the output of the fourth convolutional network is respectively connected to the input of the first sampling network and the input of the fifth splicing module; the output of the first splicing module is also connected to the input of the fourth splicing module.
[0044] In this way, the neck network presents Figure 3 the improved BiFPN structure shown in the figure, which can be denoted as the BiFPN-T structure. The BiFPN-T structure introduces a residual branch on the basis of the traditional BiFPN, retains the information of the input features through additional convolutional operations, further enhances the feature expression ability, and avoids the loss of information.
[0045] The first splicing module, the second splicing module, the third splicing module, the fourth splicing module and the fifth splicing module have the same structure, which is simply called the splicing module. The splicing module is composed of a weighted splicing unit BiCCT and a C3k2 module connected front and back.
[0046] Referring to Figure 4 , the weighted splicing unit BiCCT includes a Triplet Attention layer, a Concat-T layer, a first-dimensional stacking layer and an activation layer connected in sequence; the Triplet Attention layer adjusts the weights of multiple input feature maps; Concat-T is used to adjust the weights of the channels of the feature maps, and then combines the feature map weights and channel weights to extract dimensional features from the input feature maps for splicing to form a new feature map. The first-dimensional stacking layer stacks the feature map output by the Concat-T layer and multiple input feature maps of the BiCCT, and the stacked feature map is activated by the activation layer and then output. The activation layer can specifically adopt the Swish function.
[0047] Specifically, taking the first splicing module as an example, the input of the first splicing module is two feature maps, that is, the input of the BiCCT of the first splicing module is two feature maps, which are respectively denoted as feature map A and feature map B. The spatial dimensions of A and B are the same, both are R. At this time, assume that the weight assigned by the Triplet Attention layer to feature map A is p A , the weight assigned to feature map B is p B , p A + pB = 1; The Concat-T layer of the first splicing module is used to configure the weights of different dimensions of the feature map. Assume that the weight of the r-th dimension is denoted as q r , then q1 + q2 + … + q r + … + q R = 1; The Concat-T layer calculates the product of the feature map weight and the dimension weight as the dimension weight of the feature map, and then extracts the corresponding value with the maximum weight on each dimension to form a new feature map.
[0048] That is, let the weight of dimension r on feature map A be denoted as h A-r = p A × q r , and the weight of dimension r on feature map B be denoted as h B-r = p B × q r ;
[0049] The principle for selecting the feature of dimension r on the new feature map H is as follows: If h A-r ≥ h B-r , then the feature of feature map A on dimension r is assigned to dimension r of the new feature map H; If h A-r < h B-r , then the feature of feature map B on dimension r is assigned to dimension r of the new feature map H.
[0050] Similarly, taking the third splicing module as an example, the input of the third splicing module is three-way feature maps, that is, the input of the BiCCT of the third splicing module is three-way feature maps, denoted as feature map C, feature map D, and feature map E respectively. The spatial dimensions of C, D, and E are the same, all being R. At this time, assume that the weight assigned to feature map C by the Triplet Attention layer is p C , the weight assigned to feature map D is p D , and the weight assigned to feature map E is p E , p C + p D + p E = 1; The Concat-T layer of the third splicing module is used to configure the weights of different dimensions of the feature map. Assume that the weight of the r-th dimension is denoted as q r , then q1 + q2 + … + q r + … + q R = 1; The Concat-T layer calculates the product of the feature map weight and the dimension weight as the dimension weight of the feature map, and then extracts the corresponding value with the maximum weight on each dimension to form a new feature map.
[0051] That is, let the weight of dimension r on feature map C be denoted as h C-r = p C × q r , and the weight of dimension r on feature map D be denoted as hD-r = p D × q r ; The weight of dimension r on the feature map E is denoted as h E-r = p E × q r ;
[0052] The feature of dimension r on the new feature map H comes from the feature map F, where F ∈ {C, D, E}
[0053] The weight of dimension r on the feature map H is denoted as h H-r = MAX{h C-r , h D-r , h E-r}, which means assigning the maximum weight of the features of the feature maps C, D, and E in dimension r to dimension r of the feature map
[0054] Triplet Attention adopts a triple attention mechanism, which consists of three parallel branches. The first two branches are dedicated to capturing the cross-dimensional interactions between the channel and spatial dimensions. In the third branch, the input features are first pooled, then convolved, and finally, spatial attention weights are generated through the Sigmoid activation function. The outputs of these three branches are aggregated by taking the average to produce the final attention map. The triple attention mechanism aims to minimize the crack information loss by separately capturing the interactions in different dimensions (channel height, channel width, and spatial dimension), and then aggregating these interactions together, thus effectively balancing the computational efficiency and accuracy
[0055] In this way, Triplet Attention is introduced in BiCCT to adaptively adjust the contributions of each feature map according to the specific situation of the input features, dynamically generate feature map weights, and thus effectively improve the expression ability of low-level features. By focusing on the features of the three dimensions of channels, rows, and columns, Triplet Attention can flexibly enhance the representation of the crack region while suppressing the interference of complex backgrounds. BiCCT combines residual branches and dimension stacking to retain the information of the input features, further enhancing the expression ability of the features and avoiding information loss. The use of the Swish activation function improves the non-linear expression ability and enhances the network's adaptability to complex crack morphologies
[0056] The second improved model is denoted as YOLOv11n-BC
[0057] Refer to Figure 5 , Figure 6 , Figure 7, based on the model YOLOv11n-B, the model YOLOv11n-BC replaces the C3k2 modules in the backbone network and the C3k2 modules in each splicing module with C3k2-Faster-CAA modules.
[0058] In this embodiment, first, construct Figure 7 the Faster-CAA module shown in the figure, which includes a partial convolutional layer P Conv, a first pointwise convolutional layer, a second pointwise convolutional layer, an attention layer, and a third dimension stacking layer connected in sequence. The input of the third dimension stacking layer is also connected to the input of the Faster-CAA module; the input feature map is first used by the partial convolutional layer to extract initial features, then the initial features are strengthened by two pointwise convolutional layers, and attention features are further extracted. Then, the attention features are stacked with the input feature map in dimension and output.
[0059] To speed up the training speed and improve the accuracy of the model and avoid the problem of gradient disappearance. A batch normalization layer BN and an activation function Relu can be further set between the two pointwise convolutional layers. In this way, after the initial features are strengthened by the first pointwise convolutional layer, they are normalized and activated, and then input into the second pointwise convolutional layer for feature strengthening.
[0060] The partial convolution uses conventional convolution to extract spatial features on selected partial input channels while keeping the remaining channels unchanged. The partial convolution can more effectively reduce the amount of calculation and memory access, and reduce the computational burden on edge devices.
[0061] As Figure 6 shown, the C3k2-Faster-CAA module includes an input convolutional layer, a Split function layer, n Faster-CAA modules, a second dimension stacking layer, and an output convolutional layer connected in sequence; the second dimension stacking layer is used to stack the outputs of the Split function layer and the outputs of each Faster-CAA module in dimension. n≥1.
[0062] In this way, the Faster-CAA module introduces an attention mechanism and a residual branch, capturing the context dependence between distant pixels; through the combination of partial convolution and pointwise convolution, the feature expression ability of the central region is enhanced.
[0063] The C3k2-Faster-CAA module enhances the feature modeling channel and spatial interdependence relationship through the nesting and residual combination of multiple Faster-CAA modules, thereby improving the processing ability of crack targets under complex backgrounds and improving the detection accuracy. In subsequent specific embodiments, only one Faster-CAA module is set in the C3k2-Faster-CAA module for verification.
[0064] The improved model proposed in this embodiment can adopt PIoU, CIoU or PIoU2.
[0065] Compared with the traditional PIoU, PIoU2 (Powerful-IoU2) not only considers the position uncertainty of the crack boundary, but also adjusts the probability decay of the crack boundary through an adjustment factor to further capture the crack detection error caused by complex backgrounds. PIoU2 introduces a penalty factor P that adapts to the target size, and the penalty factor is defined as follows:
[0066] ;
[0067] In the formula, , , , is the absolute value of the distance between the corresponding edges of the predicted box and the target box, and represent the width and height of the target box; the denominator of the penalty factor P is only related to the size of the target box and has nothing to do with the size of the anchor box and the minimum circumscribed box of the target box. This design avoids the change of the penalty factor caused by the increase of the anchor box size, thus avoiding the problem of anchor box amplification; in addition, the penalty factor P will not degenerate to 0 unless the anchor box completely coincides with the target box, which ensures its continuous and effective adjustment, and P can also adapt to targets of different sizes, ensuring good performance in target detection of various sizes. The PIoU2 loss function combines the method of adaptive penalty factor and attention mechanism, effectively improving the accuracy and robustness of the target detection algorithm in complex backgrounds.
[0068] The following verifies the two improved models given by the present invention in combination with specific embodiments.
[0069] In this embodiment, the original collected dataset was taken by the author himself and contains three detection target categories: cracks in wooden components with complex backgrounds, wood knots, and water stains on the surface of wooden components. Data annotation uses the target detection annotation tool Labelimg to manually annotate the cracks in the collected images. In order to improve the performance and robustness of the network model, data augmentation methods such as random cropping, translation, rotation, brightness change, and noise addition are performed on 1200 captured photos, expanding the wooden component crack dataset to 4482 images, and dividing them according to the ratio of 8:1:1. As shown in Table 1, 3583 training images, 449 validation images, and 450 test images are obtained. Data augmentation further improves the diversity and robustness of the dataset, thereby enhancing the generalization ability of the model in complex backgrounds.
[0070] The training set is used to train the network model, the validation set is used to verify the various indicators of the trained network model, and the test set is finally input into the network model to test the image recognition effect.
[0071]
[0072] Specifically, the model training steps are as follows:
[0073] S1. Extract training samples from the training set and update the model parameters in combination with the training samples;
[0074] S2. Extract validation samples from the validation set and calculate the PIOU2 loss function of the model on the validation set;
[0075] S3. Determine whether the loss function converges; if not, update the model backward in combination with the loss function, and then return to step S1; if yes, end the training and fix the model.
[0076] In this embodiment, after the model is fixed, the detection accuracy of the model is tested on the test set. After the YOLOv11n-BC model is trained in combination with the PIOU2 loss function, it is denoted as the model YOLOv11n-CPB.
[0077] To more clearly analyze the difference in the effect of YOLOv11n-CPB and YOLOv11n in detecting wood structure cracks under complex backgrounds, as Figure 8 shown, the first row shows 5 randomly selected original images of wood component cracks under complex backgrounds, including wood cracks on different components. The second row and the third row respectively show the detection effects of the YOLOv11n model and the improved YOLOv11n-CBP model. Figure 8 Among them, YOLOv11n cannot detect some cracks, resulting in missed detections, while YOLOv11n-CBP can accurately identify all cracks under complex backgrounds. Part of the reason for this progress is that the C3k2-Faster-CAA module adaptively adjusts the response of each channel, enhancing the effect of crack feature capture ability. At the same time, during the detection process, the detection boxes of cracks in YOLOv11n are inaccurate and cannot accurately locate the cracks well, while YOLOv11n-CBP can accurately locate the cracks under complex backgrounds. The reason for this progress is that the PIoU2 loss function effectively distinguishes crack targets from complex backgrounds through a more accurate positioning mechanism, reducing the influence of poor-quality features during model training. In addition, during the detection process, YOLOv11n misidentifies the ground cracks in the environmental background as cracks in wood components. However, with the help of the BiFPN-T structure and the BiCCT module, YOLOv11n-CBP effectively integrates multi-scale feature information, integrates the detailed information of wood component cracks and other context information under complex backgrounds, and achieves accurate classification. In summary, it shows that in scenarios with more complex cracks, the detection advantage of the YOLOv11n-CBP model for wood component cracks is more obvious.
[0078] In this embodiment, YOLOv11n-CBP is further compared with various existing models shown in Table 2. The training of each model in Table 2 is carried out as shown in the above steps S1-S3. After training, all are tested on the test set, and the test results are recorded in Table 2 below.
[0079]
[0080] As can be seen from Table 2, the YOLOv11n-CBP model proposed by the present invention performs excellently in the detection of wooden structure cracks under complex backgrounds. Especially in terms of detection accuracy, the AP value exceeds other models such as YOLOv5n, YOLOv7-tiny, YOLOv8n, and YOLOv10n, reaching 86.4%. This result indicates that YOLOv11n-CBP can better detect cracks in wooden components under complex environments, especially in scenarios with more noise and more complex backgrounds.
[0081] In addition to the advantage in accuracy, YOLOv11n-CBP also has a significant advantage in computational efficiency. The trained model is only 4.2MB, and the total number of parameters is 1,827,260. Compared with other comparison models, this parameter quantity is greatly reduced, indicating that the model optimizes the computational complexity and storage requirements while maintaining high accuracy. Compared with traditional larger models, YOLOv11n-CBP can run smoothly on devices with low performance requirements, reducing the dependence on high-performance hardware, which is particularly important for crack detection in practical applications.
[0082] Although other models perform better in some indicators, the overall performance fails to exceed YOLOv11n-CBP. For example, YOLOv5s and YOLOv5x have higher recall rates, but their higher parameter quantity and computational complexity result in the need for more computational resources and storage space, which is not ideal for scenarios that require quick response and resource conservation. Although YOLOv5s is smaller than YOLOv11n-CBP in terms of parameter quantity and model size, it performs poorly in key indicators such as AP and fails to meet the high-precision requirements for crack detection.
[0083] In addition, YOLOv7 and YOLOv7-tiny fail to reach the level of YOLOv11n-CBP in indicators such as AP, Precision, and Recall. Especially in complex backgrounds, the recognition effect of the model on crack targets is limited.
[0084] Overall, YOLOv11n-CBP, with its excellent comprehensive performance and low hardware requirements, has become an ideal solution for detecting wood structure cracks in complex backgrounds. This model not only stands out in terms of detection accuracy but also performs well in terms of computational efficiency and hardware adaptability, providing strong technical support for practical applications.
[0085] The YOLOv11n-CPB model integrates the BiFPN-T structure and the C3k2_faster_CAA module on the basis of YOLOv11n and adopts the PIOU2 loss function. To further analyze the specific contributions of each improvement point to the model performance, ablation experiments were conducted on the wood component crack dataset in this embodiment, and all experiments were completed under the same parameter settings and training environment. The results of the ablation experiments are shown in Table 3 and Figure 10 as follows.
[0086]
[0087] Among them, the model structures of Experiment 1 and Experiment 4 are the traditional YOLOv11n; the model structures of Experiment 2 and Experiment 6 are those that replace all C3k2 modules with C3k2_faster_CAA modules on the basis of the traditional YOLOv11n; the model structures of Experiment 3 and Experiment 7 are the YOLOv11n-B given by the present invention; the model structures of Experiment 5 and Experiment 8 are the YOLOv11n-BC given by the present invention.
[0088] Experiments 1-8 all used the above steps S1-S3 to learn on the wood frame dataset. The PIOU2 loss function was used during the training of Experiments 1, 2, 3, and 5, and the PIOU loss function was used during the training of Experiments 4, 6, 7, and 2.
[0089] This embodiment uses the Precision-Recall (P-R) curve to provide a detailed comparison of the performance differences in detecting wood component cracks in complex backgrounds among the models in the ablation experiments. When the IoU threshold is 0.5, the P-R curves of each model are as Figure 10 follows. Among them, the YOLOv11n-CBP model shown in Experiment 8 has the most excellent detection performance in detecting wood structure cracks in complex backgrounds.
[0090] The results of Experiment 2 show that after replacing the C3k2 module with the C3k2-Faster-CAA module, the number of parameters and computational complexity of the model decreased by 8.62% and 4.76% respectively, while the model size decreased by 7.27%. This improvement was particularly prominent in the crack detection task, with Precision and AP increasing by 3.6% and 10% respectively, verifying the effect of the C3k2-Faster-CAA module in reducing redundant calculations, lowering model complexity, and enhancing the ability to capture crack features. Comparing Experiment 1 and Experiment 3, after incorporating the improved BiFPN-T into the basic model, the precision (Precision) and AP of crack detection were significantly improved. At the same time, the number of parameters and size of the model decreased by 20% and 24.21% respectively. This improvement indicates that BiFPN-T can better adapt to the diversity of crack morphologies in complex backgrounds and reduce the demand for computing resources by optimizing the feature fusion ability. Comparing Experiment 1 and Experiment 4, after introducing the PIoU2 loss function, the recall (Recall) and AP of cracks increased by 2.2% and 3.5% respectively. This shows that PIoU2 optimizes the advantage in bounding box regression through a more accurate positioning mechanism.
[0091] As can be seen from Experiment 5, after integrating the C3k2-Faster-CAA module with the BiFPN-T feature fusion network, the YOLOv11n-BC model combines the advantages of the two enhancement functions. Compared with Experiment 2, the precision (Precision) and AP of Experiment 5 increased by 4.5% and 0.5% respectively, while the model weights and the number of parameters were minimized, indicating that the model achieved a good balance between high performance and computing resources. Compared with YOLOv11n-B, the precision (Precision) and AP of YOLOv11n-BC increased by 6.3% and 7.2% respectively, further verifying the synergistic effect of these two modules on improving crack detection performance.
[0092] To further verify the performance of the model YOLOv11n-CBP. In this embodiment, the detection accuracies of YOLOv11n and YOLOv11n-CBP on cracks, knots, and water stains, as well as the average detection accuracy and recall on the three categories, were further verified, and the results are shown in Table 4 below.
[0093]
[0094] It can be seen that, compared with the original YOLOv11n model, the improved YOLOv11n-CPB has achieved significant improvements in multiple performance metrics. Specifically, the improvements of YOLOv11n-CPB in terms of precision (P), recall (R), mAP50, and mAP50-95 are 0.73%, 1.43%, 6.2%, and 13.6% respectively. In the detection of wooden component cracks in complex backgrounds, such as Figure 11 , Figure 12 as shown, the AP has increased from 70.6% to 86.4%, indicating that the average performance of the improved model at a lower IoU threshold has been significantly improved. The robustness of the model has been significantly enhanced, and it can more accurately locate and identify cracks in complex backgrounds, thus better meeting the requirements of wooden structure crack detection. In the detection tasks of knots and water stains, the APs have increased by 1.9% and 1.1% respectively. In summary, YOLOv11n-CPB maintains good detection performance within the IoU threshold range, can more clearly distinguish crack targets from background interference, and provides more reliable technical support for wooden structure crack detection.
[0095] Of course, for those skilled in the art, the present invention is not limited to the details of the above exemplary embodiments, but also includes the same or similar structures that can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claimed rights.
[0096] In addition, it should be understood that although this specification is described in terms of embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
[0097] The technologies, shapes, and structures not detailed in the present invention are all well-known technologies.
Claims
1. A training method for a wood component crack detection model under complex background based on improved YOLOv11n, characterized in that: First, the neck network of YOLOv11n is replaced. The new neck network includes the first convolutional network, the second convolutional network, the third convolutional network, the fourth convolutional network and the first sampling network, the first splicing module, the second sampling network, the second splicing module, the third splicing module, the fifth convolutional network, the fourth splicing module, the sixth convolutional network and the fifth splicing module connected in sequence; the inputs of the first, second, third and fourth convolutional networks are respectively connected to the outputs of the four network layers from shallow to deep of the backbone network; the outputs of the first, second, third and fourth convolutional networks are respectively connected to the inputs of the third, second, fourth and fifth splicing modules; the output of the second convolutional network is also connected to the input of the third splicing module, the output of the third convolutional network is also connected to the input of the first splicing module, and the output of the fourth convolutional module is also connected to the input of the first sampling network; the output of the first splicing module is also connected to the input of the fourth splicing module; then the improved model is allowed to perform machine learning on a dataset of images of wooden components with cracks marked.
2. The training method of the wood component crack detection model under complex background based on improved YOLOv11n as claimed in claim 1, characterized in that: The first splicing module, the second splicing module, the third splicing module, the fourth splicing module and the fifth splicing module have the same structure, and are all composed of weighted splicing units BiCCT and C3k2 modules connected front and back; BiCCT includes a sequentially connected Triplet Attention layer, a Concat-T layer, a first dimension superposition layer and an activation layer; the first dimension superposition layer dimensionally superimposes the output of the Concat-T layer and the input of the BiCCT.
3. The training method of the wood component crack detection model under complex background based on improved YOLOv11n as claimed in claim 2, characterized in that: Improvements to YOLOv11n also include: the C3k2 module in the splicing module is replaced by the C3k2-Faster-CAA module, which includes a sequentially connected input convolution layer, a Split function layer, one or more Faster-CAA modules, a second dimension overlay layer, and an output convolution layer; the second dimension overlay layer is used to perform dimensional overlay on the output of the Split function layer and the output of each Faster-CAA module; the Faster-CAA module introduces an attention mechanism and a residual branch to capture contextual dependencies between distant pixels.
4. The training method of the wood component crack detection model under complex background based on improved YOLOv11n as claimed in claim 3 is characterized in that: The Faster-CAA module includes a sequentially connected partial convolution layer, a first point-by-point convolution layer, a second point-by-point convolution layer, an attention layer, and a third dimension stacking layer. The input of the third dimension stacking layer is also connected to the input of the Faster-CAA module.
5. The training method of the wood component crack detection model under complex background based on improved YOLOv11n as claimed in claim 3, characterized in that: The C3k2 module in the backbone network is replaced by the C3k2-Faster-CAA module.
6. The training method for the wood component crack detection model under complex background based on improved YOLOv11n as claimed in claim 1, characterized in that: The input of the first convolutional network is connected to the 3rd layer of the backbone network, the input of the second convolutional network is connected to the 5th layer of the backbone network, the input of the third convolutional network is connected to the 7th layer of the backbone network, and the input of the fourth convolutional network is connected to the 11th layer of the backbone network.
7. The training method for a wood component crack detection model under complex background based on improved YOLOv11n according to any one of claims 1 to 6, characterized in that: The loss function used in the model training process is PIoU2 loss.
8. A method for detecting cracks in wooden components using the training method for the wooden component crack detection model under complex background based on improved YOLOv11n as claimed in any one of claims 1 to 7, characterized in that: A crack detection model is obtained by adopting the training method of the wood component crack detection model under complex background based on the improved YOLOv11n as described in any one of claims 1 to 7; taking a photo of the wood component to be detected, and inputting the photo into the crack detection model.
9. The method for detecting cracks in wooden components under complex backgrounds based on improved YOLOv11n as claimed in claim 8, characterized in that: The photos of the wood components to be inspected are enhanced and then input into the crack detection model for processing.
10. A wood component crack detection system, characterized in that: It includes a memory and a processor, the memory stores a computer program, the processor is connected to the memory, and the processor is used to execute the computer program to implement the method for detecting cracks in wooden components under complex backgrounds based on improved YOLOv11n as described in claim 8 or 9.
Citation Information
Patent Citations
Gap detection method and system based on YOLOv7 optimization and storage medium
CN117237312A
Method of using deep discriminate network model for person re-identification in image or video
US20210150268A1