Building wall defect detection method based on improved YOLOv10 network
By introducing VoVGSCSP and MCA modules into the YOLOv10 network and using MPDIoU loss function, the problems of complex background and defect diversity in the detection of building wall walls are solved, and the accuracy and efficiency of detection are improved.
Patent Information
- Application Number
- CN202411857212.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-16
AI Technical Summary
The existing YOLOv10 model has problems such as complex background interference, defect diversity, scale and position sensitivity, lighting and viewing angle changes, and insufficient data sets in the detection of building wall defects, resulting in low detection accuracy.
By replacing the C2f module of the 19th convolution layer in the YOLOv10 network as the VoVGSCSP module, and adding the MCA module in the Neck stage, a defect detection model is generated and trained using the MPDIoU loss function to improve the model's feature extraction ability and computing efficiency.
It improves the accuracy and efficiency of building wall defect detection, enhances the detection performance of the model in complex background and diverse defects, reduces the need for manual intervention, and improves work efficiency.
Smart Images

Figure CN120014427A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of object defect detection, and in particular to a building wall defect detection method based on an improved YOLOv10 network. Background Art
[0002] Building wall inspection plays a key role in building quality monitoring and maintenance, especially in ensuring the safety and durability of building structures. Traditional wall inspection methods mainly rely on manual inspection and measurement tools, which is not only time-consuming and labor-intensive, but also easily affected by human subjective factors, resulting in inaccurate inspection results. At the same time, as the complexity and scale of building structures continue to increase, traditional methods are difficult to meet the needs of efficient and accurate inspection.
[0003] In recent years, the rapid development of computer vision and deep learning technology has provided a new way to solve the above problems. Object detection methods based on convolutional neural networks (CNNs) have made significant progress in the fields of image recognition and object detection. As an end-to-end object detection method, the YOLO (You Only Look Once) series of algorithms have the advantages of strong real-time performance and high detection accuracy, and have been successful in many practical applications. As the latest version of this series of algorithms, YOLOv10 has been further optimized in terms of network structure and detection performance. However, there are still some problems in directly applying YOLOv10 to the detection of building wall defects, which requires targeted improvements.
[0004] Although YOLOv10 performs well in general object detection tasks, it still faces the following challenges in building wall defect detection tasks:
[0005] Complex background interference: Building walls usually have various textures, colors, lighting changes, and decorative elements. These complex backgrounds will interfere with the accuracy of defect detection. The existing YOLOv10 model may have difficulty accurately identifying wall defects in complex backgrounds.
[0006] Diversity of defects: There are many types of wall defects, including cracks, peeling, stains, mold, etc., and the shapes, sizes and colors of various defects vary. The existing YOLOv10 model may have insufficient detection accuracy when dealing with defects with strong diversity, especially when identifying defects of small size or irregular shape.
[0007] Scale and location sensitivity: The scale of wall defects varies greatly, from tiny cracks to large-scale peeling, and the location of defects may also be distributed in various areas of the wall. When dealing with scale and location changes, YOLOv10 may miss or misdetect, especially for small-scale defects.
[0008] Lighting and perspective changes: Building wall inspection is usually performed under different lighting conditions and perspectives, such as daytime, nighttime, sunny, cloudy, and different shooting angles. Lighting and perspective changes will affect the defect detection results. The robustness of the existing YOLOv10 model under different lighting and perspective conditions needs to be improved.
[0009] Insufficient data sets: The data sets required for building wall defect detection are usually limited, and the diversity between samples is insufficient, making it difficult to cover various practical scenarios. Directly using YOLOv10 for training may result in insufficient generalization ability of the model, which may easily lead to detection errors in practical applications.
[0010] In summary, the current YOLOv10 still has many problems when applied to building wall defect detection, resulting in low detection accuracy. Therefore, it is urgent to provide a building wall defect detection method based on an improved YOLOv10 network, which can effectively improve the detection accuracy and provide more effective technical means for building quality monitoring and maintenance. Summary of the invention
[0011] The present invention provides a building wall defect detection method based on an improved YOLOv10 network, which can improve the accuracy of building wall defect detection.
[0012] In order to achieve the above objectives, this application provides the following technical solutions:
[0013] A building wall defect detection method based on an improved YOLOv10 network comprises the following steps:
[0014] Get wall image;
[0015] Preprocessing the acquired wall image and generating an original wall image;
[0016] The C2f module of the 19th convolutional layer in the YOLOv10 network is replaced with the VoVGSCSP module, and the MCA module is added to the Neck stage of the YOLOv10 network to generate a defect detection model;
[0017] Based on the MPDIoU loss function, the defect detection model is trained;
[0018] The trained defect detection model is used to perform defect detection on wall images.
[0019] Furthermore, the processing of the defect detection model includes:
[0020] In the Backbone stage of the YOLOv10 network, feature extraction and feature fusion are performed on the original wall image to generate enhanced intermediate feature maps, enhanced high-dimensional feature maps, and multi-scale enhanced feature maps;
[0021] The enhanced intermediate feature map, enhanced high-dimensional feature map and multi-scale enhanced feature map are input into the Neck stage of the YOLOv10 network for data processing, and a low-scale feature fusion map and a high-dimensional fusion feature map are generated; in the Neck stage, the multi-scale enhanced feature map, the low-scale feature fusion map and the high-dimensional fusion feature map are further processed to generate a mid-scale feature fusion map and a high-scale feature fusion map;
[0022] The low-scale feature fusion map, the medium-scale feature fusion map, and the high-scale feature fusion map are input into the v10Detect layer in the Head stage for multi-scale target detection and generate defect detection results.
[0023] Furthermore, in the Backbone stage of the YOLOv10 network, feature extraction and feature fusion are performed on the original wall image to generate enhanced intermediate feature maps, enhanced high-dimensional feature maps, and multi-scale enhanced feature maps, including:
[0024] Input the original wall image into the initial convolutional layer for feature extraction to generate a low-level feature map;
[0025] Input the low-level feature map into the C2f_1 module for feature fusion to generate a mid-level feature map;
[0026] The intermediate feature map is input into the convolutional dimension reduction layer for feature extraction, and the intermediate feature map after feature extraction is input into the C2f_2 module for detail enhancement processing, and an enhanced intermediate feature map is generated;
[0027] The enhanced intermediate feature map is input into the SCDown_1 module for image compression processing to generate a high-dimensional feature map;
[0028] The high-dimensional feature map is input into the C2f_3 module for detail enhancement processing, and an enhanced high-dimensional feature map is generated;
[0029] The enhanced high-dimensional feature map is input into the SCDown_2 module for image compression processing, and the enhanced high-dimensional feature map after image compression processing is input into the C2f_4 module for dimensionality reduction processing, and a reduced-dimensional high-dimensional feature map is generated;
[0030] The reduced-dimensional high-dimensional feature map is input into the spatial pyramid pooling layer to extract spatial multi-scale information, and a reduced-dimensional high-dimensional feature map containing multi-scale spatial information is generated;
[0031] The reduced-dimensional high-dimensional feature map containing multi-scale spatial information is input into the adaptive attention layer for feature enhancement processing, and a multi-scale enhanced feature map is generated.
[0032] Furthermore, the enhanced intermediate feature map, enhanced high-dimensional feature map and multi-scale enhanced feature map are input into the Neck stage of the YOLOv10 network for data processing, and a low-scale feature fusion map and a high-dimensional fusion feature map are generated, including:
[0033] Input the multi-scale enhanced feature map to the first upsampling layer for upsampling;
[0034] The enhanced high-dimensional feature map and the upsampled multi-scale enhanced feature map are input into the Concat_1 module for concatenation in the channel dimension, and a high-dimensional concatenated feature map is generated;
[0035] Input the high-dimensional spliced feature map into the C2f_5 module for feature fusion and generate a high-dimensional fused feature map;
[0036] Input the high-dimensional fusion feature map into the second upsampling layer for upsampling;
[0037] The enhanced intermediate feature map and the upsampled high-dimensional fusion feature map are input into the Concat_2 module for splicing in the channel dimension, and an intermediate splicing feature map is generated;
[0038] The intermediate-level spliced feature map is input into the C2f_6 module for feature fusion, and a low-scale feature fusion map is generated.
[0039] Furthermore, in the Neck stage, the multi-scale enhanced feature map, the low-scale feature fusion map and the high-dimensional fusion feature map are further processed to generate a medium-scale feature fusion map and a high-scale feature fusion map, including:
[0040] Input the low-scale feature fusion map into the downsampling convolution layer for downsampling;
[0041] The high-dimensional fusion feature map and the downsampled low-scale feature fusion map are input into the Concat_3 module for splicing in the channel dimension, and a mid-scale splicing feature map is generated;
[0042] The mesoscale splicing feature map is input into the VoVGSCSP module for convolution fusion, and a mesoscale feature fusion map is generated;
[0043] The mesoscale feature fusion map is input into the SCDown_3 module for downsampling;
[0044] The multi-scale enhanced feature map and the downsampled mid-scale feature fusion map are input into the Concat_4 module for splicing in the channel dimension, and a high-scale spliced feature map is generated;
[0045] The high-scale spliced feature map is input into C2fCIB to introduce feature fusion of convolutional blocks;
[0046] The high-scale spliced feature map after feature fusion is input into the MCA module for weighted and feature enhancement processing in the channel dimension, and a high-scale feature fusion map is generated.
[0047] The principles and advantages of the present invention are:
[0048] First, wall images are obtained from the construction site or existing databases. These images may contain various types of defects, such as cracks, water seepage marks, peeling, etc. Then, the obtained wall images are preprocessed, including denoising, contrast enhancement, etc., to generate original wall images suitable for subsequent processing, which is conducive to improving the effect of subsequent feature extraction. The C2f module of the 19th convolutional layer in the YOLOv10 network is replaced with the VoVGSCSP module, and the MCA module is added to the Neck stage of the YOLOv10 network to generate a defect detection model, which improves the feature extraction ability and computational efficiency of the model. The defect detection model is trained based on the MPDIoU loss function; the MPDIoU loss function can more accurately measure the overlap between the predicted bounding box and the true bounding box, thereby improving the positioning accuracy of the model. Finally, the trained defect detection model is used to detect defects on the preprocessed wall images, and the model can automatically identify and mark the location and type of defects in the image.
[0049] Specifically, in the Backbone stage of the YOLOv10 network, the original wall image is feature extracted and fused to generate enhanced intermediate feature maps, enhanced high-dimensional feature maps, and multi-scale enhanced feature maps. These feature maps contain information at different levels and scales, which helps to improve the accuracy of defect detection. The low-scale feature fusion map, the medium-scale feature fusion map, and the high-scale feature fusion map are input into the v10Detect layer in the Head stage for multi-scale target detection and generate defect detection results. Among them, multi-scale detection can ensure that the model has good performance at different scales. Then, in the Backbone stage, features are extracted and enhanced through multiple C2f modules and SCDown modules to generate feature maps containing rich details. These feature maps provide a solid foundation for subsequent defect detection. In the Neck stage, multi-scale features are fused and upsampled through the Concat module and the C2f module to generate high-dimensional splicing feature maps and intermediate splicing feature maps. These fused feature maps combine information at different levels, which helps to improve the accuracy of detection. In the downsampling convolution layer and MCA module, the features are downsampled and weighted in the channel dimension to generate medium-scale feature fusion maps and high-scale feature fusion maps. These processes can further optimize the feature representation and improve the robustness of the model.
[0050] In summary, the building wall defect detection method based on the improved YOLOv10 network in this scheme is used for defect detection. By introducing the VoVGSCSP module and the MCA module, the feature extraction capability and computational efficiency of the model are enhanced, thereby improving the accuracy of defect detection. In addition, through multi-scale feature fusion and target detection, it is ensured that the model has good performance at different scales. And the entire detection process is automatically completed by the computer, reducing the need for manual intervention and improving work efficiency. Finally, this method is suitable for different types of building wall defect detection and has strong versatility and flexibility. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a schematic diagram of the architecture of a defect detection model in an embodiment of a building wall defect detection method based on an improved YOLOv10 network of the present invention.
[0052] Figure 2 The present invention is a flowchart of an embodiment of a method for detecting building wall defects based on an improved YOLOv10 network. DETAILED DESCRIPTION
[0053] The following is further described in detail through specific implementation methods:
[0054] Embodiment 1:
[0055] A building wall defect detection method based on improved YOLOv10 network, such as Figure 2 As shown, the specific steps include:
[0056] S100, acquiring a wall image; in this embodiment, a high-resolution camera or scanning equipment is used to acquire the wall image to be detected, ensuring that the image quality is clear enough for subsequent processing.
[0057] S200, preprocessing the acquired wall image and generating an original wall image; in this embodiment, the preprocessing includes image enhancement, denoising and other operations to improve the input quality of the model. In other embodiments of the present application, histogram equalization or filtering algorithms can also be used to further improve the contrast and clarity of the wall image.
[0058] S300, replace the C2f module of the 19th convolutional layer in the YOLOv10 network with the VoVGSCSP module, add the MCA module in the Neck stage of the YOLOv10 network, and generate a defect detection model.
[0059] like Figure 1 As shown, the processing process of the defect detection model includes:
[0060] In the Backbone stage of the YOLOv10 network, feature extraction and feature fusion are performed on the original wall image to generate enhanced intermediate feature maps, enhanced high-dimensional feature maps, and multi-scale enhanced feature maps; including:
[0061] The original wall image is input to the initial convolution layer through the input layer Input for feature extraction to generate a low-level feature map; in this embodiment, the initial convolution layer includes Conv_1 and Conv_2. First, Conv_1 reduces the image to 1 / 2 of the original size (P1 / 2) and generates a feature map of 64 channels; then, Conv_2 further reduces the image to 1 / 4 of the original size (P2 / 4) and generates a low-level feature map of 128 channels.
[0062] The low-level feature map is input into the C2f_1 module for feature fusion to generate an intermediate feature map; specifically, a 128-channel intermediate feature map is generated by fusing more local information.
[0063] The intermediate feature map is input into the convolutional dimension reduction layer Conv_3 for feature extraction, and the intermediate feature map after feature extraction is input into the C2f_2 module for detail enhancement processing, and an enhanced intermediate feature map is generated. Specifically, in the convolutional dimension reduction layer Conv_3, the image is further reduced to 1 / 8 of the original size (P3 / 8), and the intermediate feature map of 256 channels is extracted; then the intermediate feature map after feature extraction is input into the C2f_2 module for detail enhancement processing, enhancing the perception of crack shape and texture, and generating an enhanced intermediate feature map.
[0064] The enhanced intermediate feature map is input into the SCDown_1 module for image compression processing, and the image is further compressed to 1 / 16 of the original size to generate a 512-channel high-dimensional feature map.
[0065] The high-dimensional feature map is input into the C2f_3 module for detail enhancement and an enhanced high-dimensional feature map is generated.
[0066] The enhanced high-dimensional feature map is input into the SCDown_2 module for image compression processing, and the image is further compressed to 1 / 32 of the original size to generate a 1024-channel high-dimensional feature. The enhanced high-dimensional feature map after image compression processing is input into the C2f_4 module for dimensionality reduction processing, and a reduced-dimensional high-dimensional feature map is generated.
[0067] The reduced-dimensional high-dimensional feature map is input into the spatial pyramid pooling layer SPPF to extract spatial multi-scale information and generate a reduced-dimensional high-dimensional feature map containing multi-scale spatial information. SPPF can effectively capture receptive fields of different sizes and enhance the model's ability to detect multi-scale targets.
[0068] The reduced-dimensional high-dimensional feature map containing multi-scale spatial information is input into the adaptive attention layer PSA for feature enhancement processing, and a multi-scale enhanced feature map is generated. PSA further improves the effectiveness of feature expression by adaptively focusing on important areas in the input feature map. After PSA processing, a 1024-channel multi-scale enhanced feature map is finally generated. The multi-scale enhanced feature map combines spatial multi-scale information and the attention mechanism of important areas, providing rich information for subsequent feature fusion and target detection.
[0069] The enhanced intermediate feature map, enhanced high-dimensional feature map and multi-scale enhanced feature map are input into the Neck stage of the YOLOv10 network for data processing, and a low-scale feature fusion map and a high-dimensional fusion feature map are generated; including:
[0070] The multi-scale enhanced feature map is input into the first upsampling layer Upsample_1 for upsampling, and the spatial size of the multi-scale enhanced feature map is enlarged from the current size 1 / 32 to 1 / 16 so as to be fused with the enhanced high-dimensional feature map.
[0071] The enhanced high-dimensional feature map and the upsampled multi-scale enhanced feature map are input into the Concat_1 module for concatenation in the channel dimension, and a high-dimensional concatenated feature map with 1536 channels is generated, whose spatial size is 1 / 16. This fusion helps the model to utilize both details and semantic information.
[0072] The high-dimensional spliced feature map is input into the C2f_5 module for feature fusion, and a 512-channel high-dimensional fused feature map is generated.
[0073] The high-dimensional fused feature map is input into the second upsampling layer Upsample_2 for upsampling, and its spatial size is enlarged to 1 / 8 so as to align with the enhanced mid-level feature map.
[0074] The enhanced intermediate feature map and the upsampled high-dimensional fusion feature map are input into the Concat_2 module for splicing in the channel dimension, and a 768-channel intermediate splicing feature map suitable for detecting small targets is generated, whose spatial size is 1 / 8.
[0075] The intermediate concatenated feature map is input into the C2f_6 module for feature fusion, and a 256-channel low-scale feature fusion map is generated.
[0076] In the Neck stage, the multi-scale enhanced feature map, the low-scale feature fusion map and the high-dimensional fusion feature map are processed to generate the mid-scale feature fusion map and the high-scale feature fusion map, including:
[0077] The low-scale feature fusion map is input into the downsampling convolution layer Conv_4 for downsampling to reduce its spatial resolution to meet the detection requirements of medium-scale targets.
[0078] The high-dimensional fusion feature map and the downsampled low-scale feature fusion map are input into the Concat_3 module for splicing in the channel dimension, and a medium-scale spliced feature map with 768 channels is generated, and the spatial size is 1 / 16 of the original image. This fusion process helps to combine feature information at different levels and improve the model's detection ability for medium-sized objects.
[0079] The mid-scale splicing feature map is input into the VoVGSCSP module for convolution fusion, and a mid-scale feature fusion map is generated to better capture the details in the image, especially the complex structure. This step includes two sub-steps: multi-branch convolution processing and feature fusion.
[0080] Multi-branch convolution processing: The VoVGSCSP module processes input features through multiple parallel convolution branches. These branches have different convolution kernel sizes and can extract features of different scales, thereby improving the model's perception of multi-scale defects.
[0081] Feature fusion: The features of each branch are fused through residual connections (shortcuts) and convolutional blocks, which enhances the expressiveness of the features while retaining fine-grained spatial information.
[0082] Output: The final output is a fused and enhanced 512-channel mesoscale feature fusion map, which retains the multi-scale information of the shape and location of the cracks.
[0083] The mesoscale feature fusion map is input into the SCDown_3 module for downsampling.
[0084] The multi-scale enhanced feature map and the downsampled mid-scale feature fusion map are input into the Concat_4 module for splicing in the channel dimension, and a high-scale spliced feature map is generated to meet the detection needs of larger targets.
[0085] The high-scale spliced feature map is input into C2fCIB to introduce feature fusion of convolutional blocks.
[0086] The high-scale spliced feature map after feature fusion is input into the MCA module for weighting and feature enhancement processing in the channel dimension, and a high-scale feature fusion map is generated. Its function is to adaptively focus on the important parts of the feature map through the attention mechanism to improve the detection accuracy. The MCA module processes data through the following steps:
[0087] Channel attention: First, the MCA module weights the high-scale concatenated feature map after the input 1024-channel feature fusion in the channel dimension. The model assigns different weights to different channels according to their contributions, focusing on channels related to crack defects.
[0088] Spatial Attention: After channel attention, the MCA module also focuses on key locations in the feature map, such as the edges of cracks or salient feature areas, through an attention mechanism in the spatial dimension.
[0089] Output: Through the channel and spatial attention mechanism, the MCA module outputs more accurate feature maps, focusing on the key locations and features of crack defects. These optimized feature maps help improve the accuracy of the model's detection and classification of crack locations.
[0090] The low-scale feature fusion map, the mid-scale feature fusion map, and the high-scale feature fusion map are input to the v10Detect layer in the Head stage for multi-scale target detection and generate defect detection results. Specifically, the low-scale feature fusion map, the mid-scale feature fusion map, and the high-scale feature fusion map are input to the v10Detect_1 layer, the v10Detect_2 layer, and the v10Detect_3 layer for target detection.
[0091] S400, based on the MPDIoU loss function, trains the defect detection model. In the target detection task, ensure the reasonable division of the training set and the validation set, select appropriate hyperparameters (number of epochs and batch size), and use data enhancement techniques to improve model performance. Specifically:
[0092] First, perform data set partitioning and hyperparameter selection. Reasonable data set partitioning can avoid overfitting, that is, the model performs well on the training set, but performs poorly on the validation set or test set. In this solution, an 80 / 20 or 70 / 30 ratio is used for partitioning. Epoch (the number of times the entire training data set is traversed once) is set to 250 rounds to ensure that the model has enough time to learn the patterns in the data. Batch (Batch size determines the number of samples used each time the model weights are updated) is set to 32 to achieve a good balance between computational efficiency and stability.
[0093] Then, new training samples are generated through data augmentation to increase the diversity of data and thus improve the generalization ability of the model. The data augmentation methods used in this scheme include random cropping, rotation, flipping, and color jittering.
[0094] Finally, based on the MPDIoU loss function, the defect detection model is trained to more accurately measure the match between the predicted box and the true box.
[0095] The processing steps of MPDIoU are as follows:
[0096] enter:
[0097] A, B: Two arbitrary convex shapes, representing the predicted box and the true box respectively.
[0098] w,h: width and height of the input image.
[0099] Key point coordinates:
[0100] The coordinates of the upper left corner and lower right corner of the ground truth are (x_1^gt, y_1^gt) and (x_2^gt, y_2^gt).
[0101] The coordinates of the upper left and lower right corners of the prediction box (Prediction) are (x_1^prd, y_1^prd) and (x_2^prd, y_2^prd).
[0102] Diagonal distance calculation:
[0103] Calculate the vertex distances on the diagonal of the two boxes, i.e., d1 and d2, to measure the degree of match between the upper left corner and lower right corner of the box, respectively.
[0104] Width and height parameters:
[0105] These parameters are used to normalize coordinates or calculate relative distances between boxes.
[0106] Output:
[0107] MPDIoU: A loss that combines IoU and diagonal distance.
[0108] Therefore, through reasonable data set division, appropriate hyperparameter settings and data enhancement technology, overfitting can be effectively avoided and the generalization ability of the model can be improved. The MPDIoU loss function combines IoU and diagonal distance to more accurately measure the match between the predicted box and the real box, reducing the shape error of the box, making the loss function more robust, and providing a more accurate evaluation method, which helps the model better learn the position and shape of the bounding box.
[0109] S500 uses the trained defect detection model to perform defect detection on the wall image. Figure 1 As shown in the figure, in the last layer of the model, v10Detect, the model comprehensively analyzes feature maps of multiple scales and outputs the category and confidence of each detection box. For example, the model will identify different categories such as cracks and peelings, and give their locations and confidence in the image.
[0110] Result evaluation: As shown in Table 1, the detection results are evaluated and performance analysis is performed using indicators such as accuracy and recall to ensure that the model meets the actual application requirements.
[0111] Table 1 Model detection result evaluation table
[0112]
[0113] Armatura in vista: Exposed reinforcement, i.e. a situation where the reinforcement is exposed, usually due to inadequate covering or corrosion.
[0114] Delaminazione: Spalling is the separation between the surface and base layers of concrete.
[0115] Fessura: Cracks are the phenomenon of cracks appearing in the wall.
[0116] Spalling: Spalling, the loss of portions of concrete, usually due to material deterioration or environmental stress.
[0117] Tracce di ruggine: Rust, traces left by rust on metal components.
[0118] mAP@50: represents AP, average accuracy.
[0119] As shown in Table 2: Experiments show that the defect detection model in this scheme (yolov10n-VoVGSCSP-MCA) has an accuracy improvement of 8.3% compared with the currently popular YOLOV5, YOLOV8, YOLOV9, YOLOV10, fasterrcnn, and ssd.
[0120] Table 2 Model accuracy evaluation table
[0121] model mAP@50 Params Gflops YoloV8 64.8 3.01 8.2 YoloV9 65.0 3.24 8.7 YoloV5 64.6 7.2 16.4 ssd 58.68 26.3 62.7 Fastercnn 62.9 137.1 370.2 YoloV10 67.1 2.71 8.4 YoloV10-VoVGSCSP-MCA 71.2 2.69 8.2
[0122] In summary, in this scheme, in order to solve the problems of insufficient accuracy and low efficiency of traditional target detection methods in wall defect detection, multi-scale feature extraction, multi-channel attention mechanism and optimized loss function (MPDIoULoss) are introduced to improve the detection performance of the model under complex backgrounds. The model structure adopts VoVGSCSP module, combined with VoVNet, GhostNet and CSPNet, to further enhance the feature expression ability and reduce the computational cost. In view of the fact that the traditional YOLOV10 module may not be able to identify the complex background masking and confusion of building wall cracks in wall panel detection, the MCA (Multi-Channel Attention) attention mechanism is introduced. Experimental results show that the improved YOLOv10 model performs well in various types of wall defect detection tasks, can effectively improve the detection accuracy and speed, and provides an efficient solution for wall quality monitoring in actual projects.
[0123] The above are only embodiments of the present invention. Common knowledge such as the known specific structures and characteristics in the scheme is not described in detail here. Ordinary technicians in the relevant field know all the common technical knowledge in the technical field to which the invention belongs before the application date or priority date, can obtain all the existing technologies in the field, and have the ability to apply conventional experimental means before that date. Ordinary technicians in the relevant field can improve and implement this scheme in combination with their own abilities under the enlightenment given by this application. Some typical known structures or known methods should not become obstacles for ordinary technicians in the relevant field to implement this application. It should be pointed out that for those skilled in the art, without departing from the structure of the present invention, several deformations and improvements can be made, which should also be regarded as the scope of protection of the present invention, and these will not affect the effect of the implementation of the present invention and the practicality of the patent. The scope of protection required by this application shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.
Claims
1. A building wall defect detection method based on an improved YOLOv10 network, characterized in that: The following steps are involved: Get wall image; Preprocessing the acquired wall image and generating an original wall image; The C2f module of the 19th convolutional layer in the YOLOv10 network is replaced with the VoVGSCSP module, and the MCA module is added to the Neck stage of the YOLOv10 network to generate a defect detection model; Based on the MPDIoU loss function, the defect detection model is trained; The trained defect detection model is used to perform defect detection on wall images.
2. The building wall defect detection method based on the improved YOLOv10 network according to claim 1, characterized in that: The processing of the defect detection model includes: In the Backbone stage of the YOLOv10 network, feature extraction and feature fusion are performed on the original wall image to generate enhanced intermediate feature maps, enhanced high-dimensional feature maps, and multi-scale enhanced feature maps; The enhanced intermediate feature map, the enhanced high-dimensional feature map and the multi-scale enhanced feature map are input into the Neck stage of the YOLOv10 network for data processing, and a low-scale feature fusion map and a high-dimensional fusion feature map are generated; in the Neck stage, the multi-scale enhanced feature map, the low-scale feature fusion map and the high-dimensional fusion feature map are further processed to generate a mid-scale feature fusion map and a high-scale feature fusion map; The low-scale feature fusion map, the medium-scale feature fusion map, and the high-scale feature fusion map are input into the v10Detect layer in the Head stage for multi-scale target detection and generate defect detection results.
3. The building wall defect detection method based on the improved YOLOv10 network according to claim 2 is characterized in that: In the Backbone stage of the YOLOv10 network, feature extraction and feature fusion are performed on the original wall image to generate enhanced intermediate feature maps, enhanced high-dimensional feature maps, and multi-scale enhanced feature maps, including: Input the original wall image into the initial convolutional layer for feature extraction to generate a low-level feature map; Input the low-level feature map into the C2f_1 module for feature fusion to generate a mid-level feature map; The intermediate feature map is input into the convolutional dimension reduction layer for feature extraction, and the intermediate feature map after feature extraction is input into the C2f_2 module for detail enhancement processing, and an enhanced intermediate feature map is generated; The enhanced intermediate feature map is input into the SCDown_1 module for image compression processing to generate a high-dimensional feature map; The high-dimensional feature map is input into the C2f_3 module for detail enhancement processing, and an enhanced high-dimensional feature map is generated; The enhanced high-dimensional feature map is input into the SCDown_2 module for image compression processing, and the enhanced high-dimensional feature map after image compression processing is input into the C2f_4 module for dimensionality reduction processing, and a reduced-dimensional high-dimensional feature map is generated; The reduced-dimensional high-dimensional feature map is input into the spatial pyramid pooling layer to extract spatial multi-scale information, and a reduced-dimensional high-dimensional feature map containing multi-scale spatial information is generated; The reduced-dimensional high-dimensional feature map containing multi-scale spatial information is input into the adaptive attention layer for feature enhancement processing, and a multi-scale enhanced feature map is generated.
4. The building wall defect detection method based on the improved YOLOv10 network according to claim 2 is characterized in that: The enhanced intermediate feature map, enhanced high-dimensional feature map and multi-scale enhanced feature map are input into the Neck stage of the YOLOv10 network for data processing, and a low-scale feature fusion map and a high-dimensional fusion feature map are generated, including: Input the multi-scale enhanced feature map to the first upsampling layer for upsampling; The enhanced high-dimensional feature map and the upsampled multi-scale enhanced feature map are input into the Concat_1 module for concatenation in the channel dimension, and a high-dimensional concatenated feature map is generated; Input the high-dimensional spliced feature map into the C2f_5 module for feature fusion and generate a high-dimensional fused feature map; Input the high-dimensional fusion feature map into the second upsampling layer for upsampling; The enhanced intermediate feature map and the upsampled high-dimensional fusion feature map are input into the Concat_2 module for splicing in the channel dimension, and an intermediate splicing feature map is generated; The intermediate-level spliced feature map is input into the C2f_6 module for feature fusion, and a low-scale feature fusion map is generated.
5. The building wall defect detection method based on the improved YOLOv10 network according to claim 2 is characterized in that: In the Neck stage, the multi-scale enhanced feature map, the low-scale feature fusion map and the high-dimensional fusion feature map are processed to generate the mid-scale feature fusion map and the high-scale feature fusion map, including: Input the low-scale feature fusion map into the downsampling convolution layer for downsampling; The high-dimensional fusion feature map and the downsampled low-scale feature fusion map are input into the Concat_3 module for splicing in the channel dimension, and a mid-scale splicing feature map is generated; The mesoscale splicing feature map is input into the VoVGSCSP module for convolution fusion, and a mesoscale feature fusion map is generated; The mesoscale feature fusion map is input into the SCDown_3 module for downsampling; The multi-scale enhanced feature map and the downsampled mid-scale feature fusion map are input into the Concat_4 module for splicing in the channel dimension, and a high-scale spliced feature map is generated; The high-scale spliced feature map is input into C2fCIB to introduce feature fusion of convolutional blocks; The high-scale spliced feature map after feature fusion is input into the MCA module for weighted and feature enhancement processing in the channel dimension, and a high-scale feature fusion map is generated.
Citation Information
Patent Citations
Live pig asset checking method, system and equipment based on machine vision
CN118521252A
Weld defect detection system and method based on infrared thermal imaging and deep learning
CN118641580A
Photovoltaic cell surface defect detection method based on improved YOLOv8
CN118840353A
Cited By
Chip packaging defect detection method applied to edge device based on YOLOv11m
CN120219388A
Defect detection method based on neural network model and related device
CN121305242A
Defect detection method based on neural network model and related device
CN121305242B
Working surface detection method of wall surface polishing robot and related device
CN121353867A