Wood surface defect detection method based on multi-scale characteristics
Through multi-scale feature extraction and fusion methods, the problem of insufficient multi-scale target detection capability of the YOLO model in wood surface defect detection is solved, and the detection accuracy and robustness are improved, especially in the detection of complex backgrounds and small-scale defects.
Patent Information
- Application Number
- CN202510528549.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-09-09
AI Technical Summary
The existing YOLO model has insufficient multi-scale target detection capabilities in wood surface defect detection, especially for small-scale defects and complex backgrounds. In addition, the high-level features lack detailed information and the low-level features lack semantics, resulting in poor robustness.
A multi-scale feature extraction method is adopted, including the CBS module, CSK2 module, wavelet convolution module, spatial pyramid pooling module and channel attention mechanism, combined with the bidirectional focus-diffusion pyramid and multi-scale feature fusion module. Through feature extraction, fusion and enhancement, the model's perception ability of multi-scale defects is improved.
The model's detection accuracy and robustness for wood surface defects are improved, especially in the detection of complex backgrounds and small-scale defects, which improves detection accuracy and recall rate and reduces computational complexity.
Smart Images

Figure CN120612282A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a wood surface defect detection method based on multi-scale features. Background Art
[0002] Wood surface defects, such as knots, cracks, and decay, affect the quality and usability of wood. Traditional manual inspection methods are inefficient and susceptible to human error, making them inadequate for large-scale wood processing. Consequently, automated defect detection methods based on image processing and deep learning have become a research hotspot.
[0003] Over the past decade, with the development of computer vision and artificial intelligence technologies, wood defect detection technology has made significant progress. From initial traditional image processing methods to modern deep learning-based object detection algorithms, the detection of wood surface defects has undergone rapid evolution. Early wood defect detection methods primarily relied on traditional image processing techniques, such as edge detection, threshold segmentation, and morphological operations. While these methods can detect obvious defects in simple situations, they suffer from low accuracy when dealing with complex backgrounds or subtle defects, and struggle to cope with defects of varying types and sizes.
[0004] Currently, the YOLO (You Only Look Once) family of object detection algorithms is widely used in the image processing field due to its high efficiency and real-time detection capabilities. They have demonstrated significant advantages in areas such as video surveillance, autonomous driving, and industrial inspection. YOLO significantly improves detection speed by simultaneously performing target localization and classification in a single forward propagation. However, the YOLO algorithm still faces some challenges in wood surface defect detection. First, YOLO has certain limitations in multi-scale object detection. Wood surface defects are diverse and vary greatly in size. Traditional YOLO models are prone to missing detections or experiencing low accuracy when faced with small-scale defects such as tiny cracks and fine knots. Second, wood surface defects not only vary greatly in scale but can also be affected by factors such as texture, lighting, and background.
[0005] While existing YOLO models can capture high-level semantic features when processing multi-layer features, they fail to fully capture low-level detail features and integrate high-level features. This limits the model's ability to detect defects in complex backgrounds. In particular, the model's robustness is poor when dealing with wood surfaces with complex textures or noisy backgrounds. Summary of the Invention
[0006] In order to solve the above technical problems, the present invention provides a wood surface defect detection method based on multi-scale features, comprising the following steps:
[0007] S1. Collect different types of wood surface defect images and perform data enhancement and standardization.
[0008] S2. Build a wood surface defect detection network, input the wood surface defect image into the network, and the input image passes through the feature extraction module, feature fusion module and detection head in sequence;
[0009] S3, the feature extraction module extracts multi-scale features from the input image;
[0010] S4, the feature fusion module fuses and enhances the extracted multi-scale features, adopting a three-way parallel feature processing strategy, including upsampling branch, direct convolution branch and adaptive downsampling branch;
[0011] S5. The detection head locates and classifies defects using the feature map extracted from the feature fusion module.
[0012] The technical solution further defined in the present invention is:
[0013] Furthermore, in step S3, the input image first passes through the CBS module for convolution feature extraction; then, the feature map enters the CSK2 module for channel splitting; then, the image passes through the wavelet convolution module; next, the feature map is further processed by the CBS module; then the image passes through the spatial pyramid pooling module for multi-scale pooling; finally, the feature map passes through the channel and spatial attention mechanism module; after the above processing, the size of the feature map gradually decreases, and the final feature map sizes are 80×80, 40×40, and 20×20 respectively.
[0014] As mentioned above, in the wood surface defect detection method based on multi-scale features, in the wavelet convolution module, the feature map is first convolved and then decomposed into multiple frequency components by splitting. Each frequency component is processed by WTConv and merged into a unified feature map.
[0015] As described above, in a wood surface defect detection method based on multi-scale features, in step S3, the input feature map is first decomposed into multi-scale frequencies by discrete wavelet transform, and a decomposition filter matrix DecFilters based on discrete wavelet is constructed; the low-pass filter and high-pass filter of the given wavelet are respectively denoted as and According to the definition of wavelet convolution, four decomposition filters are constructed to extract frequency components in different directions, and the decomposition filters are repeatedly applied to each input channel:
[0016]
[0017] in, Represents the outer product operation of the tensor; the four decomposition filters represent the combination of different frequency directions, Capture low-frequency information, Extract high-frequency details in the horizontal direction, Extract high-frequency details in the vertical direction, Extract high-frequency details in the diagonal direction.
[0018] As described above, in a wood surface defect detection method based on multi-scale features, in step S3, the input feature map is subjected to wavelet decomposition using a decomposition filter to obtain four sub-bands, namely, the low-frequency component X LL And the high-frequency components X in the horizontal, vertical and diagonal directions LH 、X HL 、X HH ; After wavelet decomposition, a separate convolution operation is applied to each subband to extract the features of each frequency component:
[0019] Y LL , Y LH , Y HL , Y HH =Conv(X LL , X LH , X HL , X HH )
[0020] For the low-frequency component X LL , extract global structural information through convolution operation; for high frequency components X LH 、X HL and X HH , extract detail information through convolution, respectively Y LL 、Y LH 、Y HL and Y HH .
[0021] As described above, in a wood surface defect detection method based on multi-scale features, in step S3, after the convolution processing is completed, the convolution results of each sub-band are superimposed to obtain a comprehensive feature representation Y WT ; Finally, the original feature space is mapped by inverse wavelet transform to obtain the complete output feature map.
[0022] As described above, in a wood surface defect detection method based on multi-scale features, in step S4, the CSK2 module is first used to split the channels and perform convolution operations on each split channel; then, the feature map passes through the ASF module and is adaptively fused according to features of different scales; during the feature fusion process, a bidirectional focusing-diffusion pyramid is used to perform up and down sampling operations between high and low layers, so that the detailed features of the low layer complement the semantic features of the high layer.
[0023] As previously mentioned, a wood surface defect detection method based on multi-scale features is proposed. The ASF module adopts a three-way parallel feature processing strategy, including an upsampling branch, a direct convolution branch, and an adaptive downsampling branch. The upsampling branch restores feature resolution through 1×1 convolution and upsampling operations; the direct convolution branch maintains the original scale of the feature to extract medium-scale features; and the adaptive downsampling branch introduces more global information by expanding the receptive field.
[0024] After the features of the three branches are initially fused through the Concat operation, they enter the multi-scale feature extraction module for further processing. The multi-scale feature extraction module includes four parallel depth-wise separable convolution branches and one identity branch. The convolution kernel sizes of the four parallel depth-wise separable convolution branches are 5×5, 7×7, 9×9, and 11×11, respectively. The identity branch retains the original features through residual connection and adds them to the features extracted by convolution. The formula is as follows:
[0025]
[0026] in, Represents the multi-scale features extracted under different convolution kernel sizes, F in represents the input image, Indicates that 5×5, 7×7, 9×9, and 11×11 convolutions are performed on the image respectively;
[0027] The features after deep convolution are compressed and information reorganized through point-by-point convolution. The formula is as follows:
[0028]
[0029] Among them, F pw represents the output fusion feature, F add It represents the result of adding the input features to the multi-scale features, and PWConv represents the point-by-point convolution operation.
[0030] As described above, in a wood surface defect detection method based on multi-scale features, in step S5, during the training phase, the detection head uses 3x3 grouped normalized convolution on the feature layers of the three sizes output by the feature fusion module to adjust the number of channels of the feature map; during the feature fusion phase, the detection head introduces a shared convolution module, which includes a reparameterized 3x3 grouped normalized convolution and a 1x1 convolution. The reparameterized convolution introduces a multi-branch convolution structure during the training phase; during the inference phase, the three branch convolutions are integrated into an equivalent convolution layer using reparameterization technology.
[0031] As described above, in a wood surface defect detection method based on multi-scale features, in step S5, the detection head fuses feature maps of different scales into an equivalent convolution operation by fusing reparameterized convolution and multi-branch convolution. First, a shared convolution operation is performed on the feature maps of each scale, and reparameterized convolution is used to merge the multi-branch convolution into an equivalent convolution operation. Finally, the detection head introduces a scale adaptation module, which dynamically adjusts the prediction range of the detection box according to the size of the input object. The scale adaptation operation is expressed by the following formula:
[0032]
[0033] Output inference =W eq *X+b eq
[0034] Among them, W i represents the convolution kernel weight of the i-th branch, X represents the input feature map, b i represents the bias term of each branch; W eq =W 1×1 +W 3×3 +W avg Indicates that the 1×1 convolution W 1×1 , 3×3 convolution W 3×3 And average pooling W avg The results of the convolution operation are fused to form an equivalent convolution operation; b eq =b 1×1 +b 3×3 +b avg Indicates merging the bias terms of different branches to form the final bias term.
[0035] The beneficial effects of the present invention are:
[0036] (1) In the feature extraction module of the present invention, the diversity of training data is increased by image data enhancement to improve the generalization ability of the model; the CSK2 module is used to perform channel splitting to enhance feature diversity; the CBS module is used to enhance nonlinear expression ability; multi-scale pooling is performed by spatial pyramid pooling to enhance the perception of defects of different scales; the channel and spatial attention mechanism modules are used to focus on key areas in the image to improve the accuracy of defect detection;
[0037] (2) In this invention, a dual-path processing mechanism is designed to optimize features at different levels. On the one hand, semantics of high-level features are enhanced to improve the ability to express global information; on the other hand, richer semantic support is introduced for low-level features to improve their context perception, thereby achieving deep interaction between high-level and low-level features.
[0038] (3) In the present invention, a bidirectional focusing-diffusion pyramid is used to perform up- and down-sampling operations between high and low layers, so that the detail features of the low layer and the semantic features of the high layer can effectively complement each other, thereby improving the model's multi-scale perception ability of wood defects and effectively alleviating the problems of lack of detail information in high-layer features and insufficient semantics in low-layer features. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a schematic diagram of the overall process of the present invention;
[0040] Figure 2 Schematic diagram of the structure of a wood surface defect detection network in an embodiment of the present invention;
[0041] Figure 3 Schematic diagram of the WTConv structure in WTBlock in an embodiment of the present invention;
[0042] Figure 4 Schematic diagram of the structure of the ASF module in an embodiment of the present invention;
[0043] Figure 5 Schematic diagram of the structure of the CSCD-Head detection head in an embodiment of the present invention;
[0044] Figure 6 Schematic diagram of the visual detection results in an embodiment of the present invention. DETAILED DESCRIPTION
[0045] This embodiment provides a wood surface defect detection method based on multi-scale features. Considering the complex texture characteristics and irregular feature distribution of wood surface defects, an improved algorithm WACNet based on multi-scale features is proposed to improve the wood surface defect detection performance.
[0046] This embodiment provides a wood surface defect detection method based on multi-scale features, such as Figure 1 As shown, the following steps are included:
[0047] S1. Collect images of different types of wood surface defects, including dead knots, live knots, cracks, decay, and other defect types; perform data enhancement on the images, including rotation, flipping, cropping, and color dithering, to increase the diversity of training data and improve the generalization ability of the model; and perform image standardization to unify the image size to 512×512 to ensure that the size and color channels of the input images are consistent.
[0048] S2, build a wood surface defect detection network, and input the wood surface defect image into the network; Figure 2As shown in the figure, the wood surface defect detection network consists of three parts: feature extraction module (Backbone), feature fusion module (Neck) and detection head (Head).
[0049] S3, the feature extraction module (Backbone) is responsible for extracting multi-level features from the input image.
[0050] like Figure 2 As shown in the figure, in the Backbone part, the input image first undergoes convolution feature extraction by the CBS module, which includes 3x3 convolution, batch normalization and activation function; then, the feature map enters the CSK2 module for channel splitting to enhance feature diversity; then, the image passes through the Wavelet Convolution Block (WTBlock), in which the feature map is first convolved and then decomposed into multiple frequency components by splitting, each frequency component is processed by Wavelet Convolution (WTConv) and merged into a unified feature map; next, the feature map is further processed by the CBS module to enhance nonlinear expression capabilities; then, the image is multi-scale pooled by Spatial Pyramid Pooling (SPPF) to enhance the perception of defects at different scales; finally, the feature map is processed by the channel and spatial attention mechanism (Contextualized Cross-Scale Patch The C2PSA (C2PSA) module focuses on key areas in the image to improve defect detection accuracy. After these processing steps, the feature map size gradually decreases, and the final feature map sizes are 80×80, 40×40, and 20×20, respectively.
[0051] In this embodiment, the input feature map is first decomposed into multi-scale frequencies by discrete wavelet transform, and a decomposition filter matrix DecFilters based on discrete wavelet is constructed, as shown in FIG. Figure 3 As shown; the low-pass filter and high-pass filter of the given wavelet are respectively denoted as and According to the definition of wavelet convolution, four decomposition filters are constructed to extract frequency components in different directions. These decomposition filters are repeatedly applied to each input channel to adapt to the convolution operation of multi-channel feature maps. This generation method ensures that the decomposition filters can be directly applied to the multi-channel input feature maps of the convolutional neural network:
[0052]
[0053] in, Represents the outer product operation of the tensor; these four decomposition filters represent the combination of different frequency directions, Capture low-frequency information, Extract high-frequency details in the horizontal direction, Extract high-frequency details in the vertical direction, Extract high-frequency details in the diagonal direction.
[0054] The input feature map is decomposed by wavelet decomposition using the decomposition filter to obtain four sub-bands, which are the low-frequency component X LL And the high-frequency components X in the horizontal, vertical and diagonal directions LH 、X HL 、X HH , the low-frequency component contains the main structural information of the image, while the high-frequency component reflects the details and edge information of the image; after wavelet decomposition, an independent convolution operation is applied to each sub-band to extract the features of each frequency component:
[0055] Y LL , Y LH , Y HL , Y HH =Conv(X LL , X LH , X HL , X HH )
[0056] For the low-frequency component X LL , extract global structural information through convolution operation; for high frequency components X LH 、X HL and X HH , extract detail information through convolution, respectively Y LL 、Y LH 、Y HL and Y HH .
[0057] After the convolution process is completed, the convolution results of each sub-band are superimposed to obtain a comprehensive feature representation Y WT ; Finally, the original feature space is mapped by inverse wavelet transform to obtain the complete output feature map.
[0058] S4, the task of the feature fusion module (Neck) is to fuse and enhance the multi-scale features extracted from Backbone to help the network better identify wood defects.
[0059] In the Neck part, the CSK2 module is first used to split the channels and perform convolution operations on each split channel to enhance the feature representation of each scale; then, the feature map is processed as follows Figure 4The ASF (Adaptive Scale Fusion) module shown in the figure adaptively fuses features at different scales, further improving the network's ability to perceive defects of varying sizes. The ASF module enhances the network's flexibility and accuracy in processing information at different scales. During the fusion process, the network also utilizes a bidirectional focusing-diffusion pyramid. By performing upsampling and downsampling operations between high- and low-level layers, the network effectively complements the detailed features of the lower layers with the semantic features of the higher layers, thereby enhancing the model's multi-scale perception of wood defects.
[0060] By designing a dual-path processing mechanism to optimize features at different levels, on the one hand, the semantics of high-level features are enhanced to improve the ability to express global information; on the other hand, richer semantic support is introduced for low-level features to improve their context perception ability, thereby achieving deep interaction between high- and low-level features.
[0061] Compared with the traditional method, the bidirectional focus-diffusion pyramid structure effectively alleviates the problem of lack of detailed information in high-level features and insufficient semantics in low-level features. On this basis, this embodiment further designs the ASF (Adaptive Scale Fusion) module, such as Figure 4 As shown in Figure 3, the ASF module adopts a three-way parallel feature processing strategy, including an upsampling branch (Upsample), a direct convolution branch (Conv), and an adaptive downsampling branch (ADown).
[0062] The upsampling branch restores the feature resolution through 1×1 convolution and upsampling operations to enhance the capture of detail information; the direct convolution branch maintains the original scale of the feature to extract medium-scale features; the adaptive downsampling branch introduces more global information by expanding the receptive field.
[0063] After the three features are initially fused through the Concat operation, they enter the multi-scale feature extraction module for further processing. The multi-scale feature extraction module includes four parallel depth-wise separable convolution branches and one identity branch. The convolution kernel sizes of the four parallel depth-wise separable convolution branches are 5×5, 7×7, 9×9, and 11×11, respectively. The identity branch retains the original features through residual connections and adds them to the features extracted by convolution, achieving efficient feature fusion. The formula is as follows:
[0064]
[0065] in, Represents the multi-scale features extracted under different convolution kernel sizes, F in represents the input image, It means that 5×5, 7×7, 9×9, and 11×11 convolutions are performed on the image respectively.
[0066] The features after deep convolution are compressed and information reorganized through point-by-point convolution. The formula is as follows:
[0067]
[0068] Among them, F pw represents the output fusion feature, F add It represents the result of adding the input features to the multi-scale features, and PWConv represents the point-by-point convolution operation.
[0069] S5, the detection head (Head) is responsible for the final defect location and classification of the feature map extracted from the Neck part. The specific structure is as follows Figure 5 shown.
[0070] During the training phase, the CSCD-Head detection head uses 3x3 group normalization convolution (Group Normalization Convolution) to adjust the number of channels of the feature map for the P3, P4, and P5 feature layers of the Neck part. Through group convolution, the independence between the channels of each group is retained; in terms of feature fusion, the CSCD-Head detection head introduces a shared convolution module, which consists of a reparameterizable 3x3 group normalization convolution and a 1x1 convolution. The reparameterizable convolution introduces a multi-branch convolution structure during the training phase to capture feature information of different scales and levels.
[0071] During the inference phase, the reparameterization technology integrates the three branch convolutions (3x3, 5x5, and 7x7 convolutions) into an equivalent convolution layer to reduce the computational complexity of the inference phase, thereby ensuring efficient inference. The design of the shared convolution module enables the features of the three scales P3, P4, and P5 to be efficiently fused in the shared convolution, reducing repeated calculations and improving the model's ability to express features of different scales.
[0072] In order to improve the efficiency of multi-scale object detection and reduce redundant computation, a reparameterizable shared convolution detection head is designed. This detection head fuses feature maps of different scales into an equivalent convolution operation by fusing reparameterized convolution and multi-branch convolution, thereby reducing the computational complexity.
[0073] like Figure 5 As shown in the figure, first, a shared convolution operation is performed on the feature maps of each scale, and a reparameterized convolution is used to merge the multi-branch convolution into an equivalent convolution operation, thereby further reducing the amount of computation during inference. Finally, in order to adapt to targets of different sizes, the detection head introduces a scale adaptation module, which dynamically adjusts the prediction range of the detection box according to the size of the input target. The scale adaptation operation can be expressed by the following formula:
[0074]
[0075] Output inference =W eq *X+b eq
[0076] Among them, W i represents the convolution kernel weight of the i-th branch, X represents the input feature map, b i represents the bias term of each branch; W eq =W 1×1 +W 3×3 +W avg Indicates that the 1×1 convolution W 1×1 , 3×3 convolution W 3×3 And average pooling W avg The results of the convolution operation are fused to form an equivalent convolution operation; b eq =b 1×1 +b 3×3 +b avg Indicates merging the bias terms of different branches to form the final bias term.
[0077] In the experiment, YOLOv11 was used as the backbone network and combined with a multi-scale feature extraction module to optimize the detection effect. The Adam optimizer was used in the model training process. The batch size was 32, and the mean square error (MSE) loss function was used as the optimization objective to minimize the error of the detection box and the classification error.
[0078] The experimental environment is the Windows operating system, and the NVIDIA GeForce RTX 4090 graphics card is used for training. CUDA 11.6 and cuDNN 8.0 are used to accelerate the GPU training process. All models are implemented in the PyTorch 2.1.0 framework and developed with Python 3.10.
[0079] To evaluate the performance of the method in this embodiment in wood surface defect detection, a wood defect dataset was used and compared with various existing object detection methods. The experimental results were evaluated using indicators such as mAP@50, precision, and recall.
[0080] As shown in Table 1 below, compared to the existing YOLOv11 model, this embodiment demonstrates significant gains and improvements in detection accuracy in multiple aspects. Compared with the existing YOLOv11 model, the model of this embodiment improves the mAP@50 indicator by 2.3 percentage points, the precision by 4.0 percentage points, and the recall by 2.7 percentage points. These improvements demonstrate that this embodiment excels in detection accuracy, particularly in the detection of complex backgrounds and small-scale defects, significantly outperforming YOLOv11 in scenarios.
[0081] Table 1 Comparison of the accuracy of various algorithms
[0082]
[0083] As shown in Table 2 below, the average detection accuracy of the baseline model is 93.3%. By adding the WTBlock module to the baseline model and introducing multi-scale deep convolution to enhance feature extraction, the average detection accuracy of the model is improved by 0.5%, and the model complexity is slightly reduced.
[0084] When the CSCD module was introduced, the average detection accuracy of the model increased by 1.1%, and the recall rate also increased significantly, while the model complexity and computational complexity were reduced. After the AFEM module was introduced, the average detection accuracy increased by 1.4%, the recall rate also increased significantly, and the computational complexity of the model increased slightly.
[0085] After combining the WTBlock and CSCD modules, the model improves the average detection accuracy by 1.1% while maintaining a low computational complexity, while the recall rate decreases slightly; when combining the WTBlock and AFEM modules, the average detection accuracy increases by 1.6%, and the feature extraction effect is better, but the computational complexity increases; when combining the AFEM and CSCD modules, the recall rate performs better, and the average detection accuracy increases to 0.653.
[0086] Finally, by adding the WTBlock, AFEM, and CSCD modules simultaneously, the model achieved the best performance, with an average detection accuracy of 95.6%, an improvement of 6.3% compared to the basic model.
[0087] Table 2 Ablation experiment
[0088]
[0089] By comparing different categories, the detection results of the wood defect detection model proposed in this embodiment on the self-made dataset are as follows: Figure 6 As shown in the figure, the model can accurately identify live knots, dead knots, wavy defect cracks, edge knots and small knots, and the confidence values displayed in the detection box are all high (mostly between 0.80-0.90), indicating that the model has good detection accuracy and stability.
[0090] Among them, the detection performance of complex texture defects such as cracks and corrugation defects, as well as detailed targets such as sections, is particularly outstanding, accurately selecting the targets and giving high confidence; in addition, the model is also stable when processing edge targets and multi-target scenarios, and can achieve effective detection at different feature levels; the results show that the model of this embodiment has significant advantages in multi-scale feature extraction and contextual information capture, and can achieve high-precision target detection in complex backgrounds and various defects, without obvious false detections or missed detections.
[0091] In addition to the above embodiments, the present invention may also have other implementations. Any technical solution formed by equivalent replacement or equivalent transformation falls within the scope of protection required by the present invention.
Claims
1. A wood surface defect detection method based on multi-scale features, characterized by: The following steps are involved: S1. Collect different types of wood surface defect images and perform data enhancement and standardization. S2. Build a wood surface defect detection network, input the wood surface defect image into the network, and the input image passes through the feature extraction module, feature fusion module and detection head in sequence; S3, the feature extraction module extracts multi-scale features from the input image; S4, the feature fusion module fuses and enhances the extracted multi-scale features, adopting a three-way parallel feature processing strategy, including upsampling branch, direct convolution branch and adaptive downsampling branch; S5. The detection head locates and classifies defects using the feature map extracted from the feature fusion module.
2. The method for detecting wood surface defects based on multi-scale features according to claim 1, characterized in that: In step S3, the input image first passes through the CBS module for convolution feature extraction; then, the feature map enters the CSK2 module for channel splitting; then, the image passes through the wavelet convolution module; next, the feature map is further processed by the CBS module; then, the image passes through the spatial pyramid pooling module for multi-scale pooling; finally, the feature map passes through the channel and spatial attention mechanism module; after the above processing, the size of the feature map gradually decreases, and the final feature map sizes are 80×80, 40×40, and 20×20 respectively.
3. The method for detecting wood surface defects based on multi-scale features according to claim 2, characterized in that: In the wavelet convolution module, the feature map is first subjected to a convolution operation and then decomposed into multiple frequency components by splitting. Each frequency component is processed by WTConv and merged into a unified feature map.
4. The method for detecting wood surface defects based on multi-scale features according to claim 2, wherein: In step S3, the input feature map is first subjected to multi-scale frequency decomposition by discrete wavelet transform, and a decomposition filter matrix DecFilters based on discrete wavelet is constructed; The low-pass filter and high-pass filter of the given wavelet are denoted as and According to the definition of wavelet convolution, four decomposition filters are constructed to extract frequency components in different directions, and the decomposition filters are repeatedly applied to each input channel: in, Represents the outer product operation of the tensor; the four decomposition filters represent the combination of different frequency directions, Capture low-frequency information, Extract high-frequency details in the horizontal direction, Extract high-frequency details in the vertical direction, Extract high-frequency details in the diagonal direction.
5. The method for detecting wood surface defects based on multi-scale features according to claim 4, characterized in that: In step S3, the input feature map is subjected to wavelet decomposition using a decomposition filter to obtain four sub-bands, which are low-frequency components X LL And the high-frequency components X in the horizontal, vertical and diagonal directions LH 、X HL 、X HH ; After wavelet decomposition, a separate convolution operation is applied to each subband to extract the features of each frequency component: AND LL ,AND LH ,AND HL ,AND HH =Conv(X LL ,X LH ,X HL ,X HH ) For the low-frequency component X LL , extract global structural information through convolution operation; for high frequency components X LH 、X HL and X HH , extract detail information through convolution, respectively Y LL 、Y LH 、Y HL and Y HH .
6. The method for detecting wood surface defects based on multi-scale features according to claim 5, characterized in that: In step S3, after the convolution process is completed, the convolution results of each sub-band are superimposed to obtain a comprehensive feature representation Y WT ; Finally, the original feature space is mapped by inverse wavelet transform to obtain the complete output feature map.
7. The method for detecting wood surface defects based on multi-scale features according to claim 1, characterized in that: In step S4, the CSK2 module is first used to split the channels and perform convolution operations on each split channel; then, the feature map is passed through the ASF module to perform adaptive fusion based on features of different scales; During the feature fusion process, a bidirectional focusing-diffusion pyramid is used to perform up- and down-sampling operations between high and low layers, so that the detail features of the low layers complement the semantic features of the high layers.
8. The method for detecting wood surface defects based on multi-scale features according to claim 7, characterized in that: The ASF module adopts a three-way parallel feature processing strategy, including an upsampling branch, a direct convolution branch, and an adaptive downsampling branch. The upsampling branch restores feature resolution through 1×1 convolution and upsampling operations. The direct convolution branch maintains the original scale of the features to extract medium-scale features; The adaptive downsampling branch introduces more global information by expanding the receptive field; After the features of the three branches are initially fused through the Concat operation, they enter the multi-scale feature extraction module for further processing. The multi-scale feature extraction module includes four parallel depth-wise separable convolution branches and one identity branch. The convolution kernel sizes of the four parallel depth-wise separable convolution branches are 5×5, 7×7, 9×9, and 11×11, respectively. The identity branch retains the original features through residual connection and adds them to the features extracted by convolution. The formula is as follows: in, Represents the multi-scale features extracted under different convolution kernel sizes, F in represents the input image, Indicates that 5×5, 7×7, 9×9, and 11×11 convolutions are performed on the image respectively; The features after deep convolution are compressed and information reorganized through point-by-point convolution. The formula is as follows: Among them, F pw represents the output fusion feature, F add It represents the result of adding the input features to the multi-scale features, and PWConv represents the point-by-point convolution operation.
9. The method for detecting wood surface defects based on multi-scale features according to claim 1, characterized in that: In step S5, during the training phase, the detection head uses 3x3 group normalized convolution on the feature layers of the three sizes output by the feature fusion module to adjust the number of channels of the feature map; In the feature fusion stage, the detection head introduces a shared convolution module, which includes re-parameterized 3x3 group normalized convolution and 1x1 convolution. The re-parameterized convolution introduces a multi-branch convolution structure in the training stage; in the inference stage, the re-parameterization technology is used to integrate the three branch convolutions into an equivalent convolution layer.
10. The method for detecting wood surface defects based on multi-scale features according to claim 9, characterized in that: In step S5, the detection head fuses feature maps of different scales into an equivalent convolution operation by fusing reparameterized convolution and multi-branch convolution. First, a shared convolution operation is performed on the feature map of each scale, and the multi-branch convolution is merged into an equivalent convolution operation using reparameterized convolution. Finally, the detection head introduces a scale adaptive module, which dynamically adjusts the prediction range of the detection box according to the size of the input object. The scale adaptive operation is expressed by the following formula: Output inference =W eq *X+b eq Among them, W i represents the convolution kernel weight of the i-th branch, X represents the input feature map, b i represents the bias term of each branch; W eq =W 1×1 +W 3×3 +W avg Indicates that the 1×1 convolution W 1×1 , 3×3 convolution W 3×3 And average pooling W avg The results of the convolution operation are fused to form an equivalent convolution operation; b eq =b 1×1 +b 3×3 +b avg Indicates merging the bias terms of different branches to form the final bias term.
Citation Information
Cited By
Image edge fluctuation feature intelligent detection method
CN120976566A
Underwater target detection encoder, detection system and method
CN122024028A