Solid wood floor classification and identification method based on asymmetric multi-branch convolution
By using technical means such as asymmetric multi-branch convolution module and hybrid attention module in the classification and identification of solid wood floors, the problem of low accuracy and efficiency of solid wood floor texture and color classification in the existing technology is solved, and a more efficient and accurate solid wood floor image classification is achieved.
Patent Information
- Application Number
- CN202510225904.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art has problems with low accuracy and efficiency in the texture and color classification of solid wood floors, especially the classification of color aberration wooden floors and white-edged wooden floors is less researched.
A solid wood floor classification recognition method based on asymmetric multi-branch convolution is adopted, and the comprehensive classification results of floor color and texture are output through steps such as edge detection, downsampling, LayerNorm layer, CAFNet Block module and global average pooling layer, combined with asymmetric multi-branch convolution module, hybrid attention module and regularization module.
It significantly improves the accuracy and efficiency of image classification of solid wood floors, especially when identifying categories such as white-edged wooden boards and color-difference wooden boards, which improves the model's perception of textures in different directions and the ability to excavate and identify key areas of solid wood floors.
Smart Images

Figure CN120147729A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of classification of solid wood floors, and particularly relates to a method for classifying and identifying solid wood floors based on asymmetric multi-branch convolution. Background Art
[0002] In the furniture industry, the uniformity of the surface color of solid wood floors directly affects the practicality and aesthetics of solid wood floors. However, due to multiple factors such as the growth years of wood, environmental differences, and floor processing technology, color differences or white edge problems may occur on the surface of solid wood floors. Since neural networks provide simple and efficient detection methods, various high-performance networks have been used for texture recognition and color classification in images, and good results have been achieved. However, although these methods perform well on their respective datasets, most of them only focus on texture recognition or color classification of solid wood floors, lacking a method for simultaneously recognizing the texture and color of solid wood floors and having low classification accuracy and efficiency. In addition, in the color classification task of solid wood floors, there is a lack of classification research on color-differentiated wood floors and white-edge wood floors. Summary of the Invention
[0003] Object of the Invention: Aiming at the above problems existing in the prior art, the present invention provides a method for classifying and identifying the color and texture of solid wood floors.
[0004] Technical Solution: A method for classifying and identifying solid wood floors based on asymmetric multi-branch convolution of the present invention specifically includes the following steps:
[0005] S1: Perform edge detection on the collected floor color image, remove the background color, and perform size scaling to obtain a standardized floor image F1;
[0006] S2: Perform downsampling processing on the standardized floor image F1 and then pass through the LayerNorm layer to obtain a feature map F2;
[0007] S3: Input the feature map F2 into the CAFNet Block module and stack it 3 times to obtain a feature map F3;
[0008] S4: Input the feature map F3 into the downsampling module and then pass through the CAFNet Block module stacked 3 times to obtain a feature map F4;
[0009] S5: Input the feature map F4 into the downsampling module and then pass through the CAFNet Block module stacked 9 times to obtain a feature map F5;
[0010] S6: Input the feature map F5 into the downsampling module and then pass through the CAFNet Block module stacked 3 times to obtain a feature map F6;
[0011] S7: Output the comprehensive classification results of the floor color and texture by passing the feature map F6 through a global average pooling layer, a Layer Norm layer, and a fully connected layer;
[0012] Among them, the CAFNet Block module includes an asymmetric multi-branch convolution module, a hybrid attention module, and a regularization module.
[0013] Preferably, the longitudinal field of view of the collected wooden floor image is at least 1.2 times the width of the wooden floor to prevent incomplete display of the solid wood floor part in the image taken at an inclined angle. The number of pixels in the height of the image is h, the number of pixels in the width is w, where w is at least 4 times h, and the number of image channels is 3, namely the R, G, and B channels.
[0014] Preferably, the asymmetric multi-branch convolution module (ACM) divides the input feature map X m×n×k into 3 groups of parallel branches (X m×n×w , X m×n×h , X m×n×id ) along the channel dimension k, where X m×n×w is the feature map width branch, X m×n×h is the feature map height branch, X m×n×id is the identity mapping branch, and the number of channels of each branch is g, that is, w = h = id = g; then a sliding window mechanism is used to extract features from each obtained feature map branch, that is
[0015]
[0016] X’ id = X m×n×id
[0017] where DWConv is the depthwise separable convolution operation, Concat is the concatenation operation, X' h is the output of the height-direction asymmetric convolution branch, X' w is the output of the width-direction asymmetric convolution branch, '' id is the identity mapping branch, 1×k w , k h ×1 is the asymmetric convolution kernel size, and the height and width of the feature maps output by each branch are the same as those of the input feature map; finally, the outputs of each branch are connected and merged, that is
[0018] X′ = Concat(X′ w , X′ h , X′ id )
[0019] where X' is the result after concatenating the outputs of different branches.
[0020] Preferably, the height and width of the feature map output by the asymmetric convolution branch in the asymmetric multi-branch convolution module (ACM) are calculated by the following formula:
[0021]
[0022] Where W out is the width of the output feature map, H out is the height of the output feature map, W is the width of the input feature map, H is the height of the input feature map, k w is the width of the asymmetric convolution kernel, k h is the height of the asymmetric convolution kernel, P w is the unilateral padding value in the width direction, P h is the unilateral padding value in the height direction. Since the size of the input feature map is the same as that of the output feature map, that is, W out = W, H out = H, calculate the values of P w and P h values.
[0023] Preferably, the hybrid attention module (MAM) first upsamples the feature map extracted by the asymmetric multi-branch convolution module (ACM) through a 1×1 convolution kernel, passes through the GELU activation function, then downsamples through a 1×1 convolution kernel, and then passes through the CBAM attention module, where the CBAM attention module includes a channel attention module (CAM) and a spatial attention module (SAM).
[0024] Preferably, the channel attention module (CAM) first performs global max pooling and global average pooling on the input feature map respectively to obtain two 1×1×C feature vectors, then sends them into a multi-layer perceptron (MLP) respectively, and then adds and sigmoid-activates the two feature vectors output by the multi-layer perceptron (MLP). The generated feature vector is the attention weight of the channel dimension of the original input feature map. Multiply the obtained channel dimension attention weight by the feature map to be processed to obtain the output feature map of the channel attention module.
[0025] Preferably, the spatial attention module (SAM) takes the output feature map of the channel attention module as the input feature map. First, it performs global max pooling and global average pooling on the input feature map to obtain two H×W×1 feature maps, concatenates them in the channel dimension, then uses a convolution kernel with a size of 7×7 and a stride of 1 for feature extraction, reduces the dimension to 1, and then generates a spatial attention weight through the sigmod function, and multiplies it with the input feature map of this module to regenerate the feature map.
[0026] Preferably, the regularization module (RM) is composed of a Layer Scale layer and a Drop Path layer connected in series.
[0027] Preferably, the loss function used for the parameters of each layer and each module during training is Focal Loss Pro, and the calculation formula is:
[0028] FL(p t ) = -α t (1 - p t )γ log(p t ) - β 1 log(p HS ) - β 2 log(p HC )
[0029] where p t is the prediction probability of the model for the target class, α t is the weight of the samples of the target class, γ is the focal factor, β 1 is the weight for increasing the error of the patterned dark solid wood floor, β 2 is the weight for increasing the error of the patterned color-differentiated wood board, p HS is the prediction probability of the patterned dark solid wood floor, p HC is the prediction probability of the patterned color-differentiated solid wood floor.
[0030] Preferably, the comprehensive classification results of the floor color and texture in S7 include 8 categories: patterned white-edge wood board, patterned color-differentiated wood board, patterned light-colored wood board, patterned dark-colored wood board, straight-grained white-edge wood board, straight-grained color-differentiated wood board, straight-grained light-colored wood board, and straight-grained dark-colored wood board.
[0031] Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0032] (1) The present invention uses the Canny edge detection algorithm to remove the black background to reduce the data volume;
[0033] (2) The asymmetric multi-branch convolution module (ACM) adopted by the present invention can flexibly capture the directional changes of features, efficiently extract horizontal and vertical features in the image, and significantly enhance the model's perception ability of textures in different directions while reducing the computational amount;
[0034] (3) The hybrid attention module (MAM) adopted by the present invention, two 1×1 convolutions promote information exchange between channels by scaling the channels, and the CBAM module enables the convolution process to perceive the spatial position information and channel information of the feature map, enhancing the mining and recognition of key regions and potential key features of the solid wood floor;
[0035] (4) The Focal Loss Pro function reduces the impact of uneven sample distribution by increasing the weights of samples with a smaller quantity, enabling the model to pay more attention to samples that are difficult to classify. Additionally, increasing the loss for samples of specific categories causes the model to focus on predicting samples of specific categories, thereby improving the overall accuracy of solid wood floor image classification. Description of the Drawings
[0036] Figure 1 Schematic diagram of the overall architecture of the present invention;
[0037] Figure 2 Overall module diagram of CBAM of the present invention;
[0038] Figure 3 Channel attention module diagram in the CBAM module of the present invention;
[0039] Figure 4 Spatial attention module diagram in the CBAM module of the present invention;
[0040] Figure 5 Schematic diagram for comparing the confusion matrices of the present invention with different models. In the figure, (a) is the AlexNet model; (b) is the VGG model; (c) is the ConvNeXt model; (d) is the network model of the present invention; Figure 5 In the confusion matrix diagram, HB is the white-edge wooden board with patterns, HC is the wooden board with pattern color difference, HQ is the light-colored wooden board with patterns, HS is the dark-colored wooden board with patterns, ZB is the straight-grained white-edge wooden board, ZC is the straight-grained wooden board with color difference, ZQ is the light-colored straight-grained wooden board, and ZS is the dark-colored straight-grained wooden board. Detailed Implementation Manner
[0041] This embodiment provides a method for classifying and recognizing solid wood floors based on asymmetric multi-branch convolution. The overall structure is as Figure 1 shown, including a downsampling layer, a Layer Norm layer, a downsampling module, a CAFNet Block module, a global average pooling layer, and a fully connected layer for classification. Among them, the downsampling module consists of a Layer Norm layer and a downsampling layer connected in series. The CAFNet Block module includes an asymmetric multi-branch convolution module (ACM), a hybrid attention module (MAM), and a regularization module (RM), and specifically includes the following steps:
[0042] S1: Use the Candy algorithm to perform edge detection on the collected floor color images, remove the background color, and perform size scaling to obtain a standardized image F1. The height of the scaled image is 224 pixels, and the width is 896 pixels.
[0043] S2: Apply a regular convolution with a convolution kernel tensor of shape 4×4×3 and a stride of 4 to the image processed in S1 to form a feature map of size 56×224×96, and then obtain the feature map F2 through the Layer Norm layer;
[0044] S3: Input the feature map F2 processed in S2 into the CAFNet Block module stacked three times to obtain the feature map F3; the feature map F2 passes through the Asymmetric Multi-branch Convolution Module (ACM), the Hybrid Attention Module (MAM), and the Regularization Module (RM) in sequence in the CAFNet Block module.
[0045] S4: Input F3 into the downsampling module and then pass it through the CAFNet Block module stacked 3 times to obtain a feature map F4 of size 56×224×96;
[0046] S5: Input F4 into the downsampling module and then pass it through the CAFNet Block module stacked 9 times to obtain a feature map F5 of size 28×112×192;
[0047] S6: Input F5 into the downsampling module and then pass it through the CAFNet Block module stacked 3 times to obtain a feature map F6 of size 14×566×384;
[0048] S7: After passing F6 through the global average pooling layer, the Layer Norm layer, and the fully connected layer, obtain the comprehensive classification result of the floor color and texture.
[0049] The longitudinal field of view of the collected wooden floor image is at least 1.2 times the width of the wooden floor to prevent incomplete display of the solid wood floor part in the image taken at an inclined angle. The image has a height pixel count of 1024, a width pixel count of 4096, and 3 image channels, namely the R, G, and B channels.
[0050] The Asymmetric Multi-branch Convolution Module (ACM) divides the input feature map X m×n×k along the channel dimension k into 3 groups of parallel branches (X m×n×w , X m×n×h , X m×n×id ), where X m×n×w is the feature map width branch, X m×n×h is the feature map height branch, X m×n×id is the identity mapping branch, and the number of channels of each branch is g, that is, w = h = id = g; then adopt the sliding window mechanism to perform feature extraction on the obtained feature map branches, that is
[0051]
[0052] X′ id = X m×n×id
[0053] In the formula, DWConv is the depthwise separable convolution operation, Concat is the concatenation operation, and X' h is the output of the height-direction asymmetric convolution branch, and X' w is the output of the width-direction asymmetric convolution branch, and X' id is the identity mapping branch, 1×k w , k h ×1 is the size of the asymmetric convolution kernel. Considering the size of the receptive field and the required computational amount, depthwise separable convolutions with asymmetric convolution kernel tensor shapes of 1×11×g and 3×1×g are adopted. The height and width of the feature maps output by each branch are the same as those of the input feature map; finally, the outputs of each branch are connected and merged, that is
[0054] X′ = Concat(X′ w , X′ h , X′ id )
[0055] In the formula, X' is the result after concatenating the outputs of different branches.
[0056] For the asymmetric multi-branch convolution module (ACM), the height and width of the feature maps output by the asymmetric convolution branch are calculated by the following formula:
[0057]
[0058] where W out is the width of the output feature map, H out is the height of the output feature map, W is the width of the input feature map, H is the height of the input feature map, k w is the width of the asymmetric convolution kernel, k h is the height of the asymmetric convolution kernel, P w is the unilateral padding value in the width direction, and P h is the unilateral padding value in the height direction. Since the size of the input feature map is the same as that of the output feature map, that is, W out = W, H out = H, the values of P w and P h are calculated. By adopting an asymmetric convolution kernel tensor shape of 1×11×g with a stride of 1, the corresponding P w is calculated to be 5, and P h is 0. By an asymmetric convolution kernel tensor shape of 3×1×g with a stride of 1, the corresponding P w is calculated to be 0, and P h is 1.
[0059] The Mixed Attention Module (MAM) first upsamples the feature map extracted by the Asymmetric Multi-branch Convolution Module (ACM) using a 1×1 convolutional kernel to 4 times the number of channels of the original feature map, applies the GELU activation function, and then downsamples it back to the number of channels of the original feature map using a 1×1 convolutional kernel to enhance the information exchange between the channels of the feature map. Then, it passes through the CBAM attention module, where the CBAM attention module includes a Channel Attention Module (CAM) and a Spatial Attention Module (SAM).
[0060] The Channel Attention Module (CAM) first performs global max pooling and global average pooling on the input feature map respectively to obtain two 1×1×C feature vectors, then sends them into a Multi-Layer Perceptron (MLP) respectively. After that, the two feature vectors output by the Multi-Layer Perceptron (MLP) are added and activated by Sigmoid. The generated feature vector is the attention weight of the channel dimension of the original input feature map. Multiply the obtained attention weight of the channel dimension with the feature map to be processed to obtain the output feature map of the channel attention module.
[0061] The Spatial Attention Module (SAM) takes the output feature map of the channel attention module as the input feature map. First, it performs global max pooling and global average pooling on the input feature map to obtain two H×W×1 feature maps, concatenates them in the channel dimension, then uses a convolutional kernel with a size of 7×7 and a stride of 1 for feature extraction, reduces the dimension to 1, and then generates the spatial attention weight through the Sigmod function, and multiplies it with the input feature map of this module to regenerate the feature map.
[0062] The Regularization Module (RM) is composed of a Layer Scale layer and a Drop Path layer in series.
[0063] The loss function used for the parameters of each layer and each module during training is Focal Loss Pro, and the calculation formula is:
[0064] FL(p t )=-α t (1-p t ) γ log(p t )-β 1 log(p HS )-β 2 log(p HC )
[0065] where p t为 is the predicted probability of the model for the target class, α t is the weight of the samples of the target class, γ is the focal factor, β 1 is the weight for increasing the error of the patterned dark solid wood floor, β 2Increase the weight of the error for the wooden board with pattern color difference, p HS Predicted probability for the solid wood floor with dark pattern, p HC Predicted probability for the solid wood floor with pattern color difference, take β 1 = 0.01, β 2 = 0.01, α t = 0.25, γ = 2.
[0066] The comprehensive classification results of the floor color and texture in S7 include 8 categories: wooden board with white pattern edge, wooden board with pattern color difference, wooden board with light pattern, wooden board with dark pattern, straight-grained wooden board with white edge, straight-grained wooden board with color difference, straight-grained wooden board with light color, and straight-grained wooden board with dark color.
[0067] Using the above solid wood floor classification and recognition method, taking the collected wooden floor image data as the data set, verifying the recognition performance of the solid wood floor classification and recognition method based on asymmetric multi-branch convolution proposed by the present invention, and comparing with the existing methods, the results are shown in Table 1.
[0068] Model Accuracy Precision Recall F1 Score Number of Parameters AlexNet 83.8% 85.0% 83.9% 84.5% <![CDATA[4.09×10 6 > VGG16 87.6% 88.3% 87.6% 88.0% <![CDATA[13.8×10 7 > ConvNeXt 92.4% 93.8% 92.2% 93.0% <![CDATA[2.90×10 7 > The method of the present invention 94.3% 94.6% 94.2% 94.3% <![CDATA[2.76×10 7 >
[0069] In addition, in order to verify the superiority of the model of the present invention relative to the classical classification models (AlexNet, VGG16, ConvNeXt), a confusion matrix is introduced to evaluate the classification results of the solid wood floor images, as shown in Figure 5 , and through comparative analysis Figure 5 It can be seen that compared with the other 3 classical classification models, the recognition accuracy of the model of the present invention for the wooden board with white pattern edge, wooden board with pattern color difference, wooden board with light pattern, wooden board with dark pattern, straight-grained wooden board with white edge, straight-grained wooden board with color difference, straight-grained wooden board with light color, and straight-grained wooden board with dark color has all been improved, and the recognition accuracy rate can reach more than 84.0%, which are 96.2%, 91.6%, 99.7%, 92.6%, 99.3%, 92.2%, 99.7% and 86.1% respectively.
[0070] The specific implementation manner is only a preferred embodiment of the present invention, and is not used to limit the implementation and the scope of the claims of the present invention. Any equivalent changes and modifications made according to the content of the patent protection scope of the present invention shall be included in the scope of the patent application of the present invention.
Claims
1. A solid wood floor classification and recognition method based on asymmetric multi-branch convolution, characterized in that: Follow these steps: S1: Perform edge detection on the collected floor color image, remove the background color, and resize it to obtain a standardized floor image F1; S2: downsampling the standardized floor image F1 and passing it through the LayerNorm layer to obtain a feature map F2; S3: Input the feature map F2 into the CAFNet Block module and stack it 3 times to obtain the feature map F3; S4: The feature map F3 is input into the downsampling module and then stacked three times through the CAFNet Block module to obtain the feature map F4; S5: The feature map F4 is input into the downsampling module and then stacked 9 times through the CAFNet Block module to obtain the feature map F5; S6: The feature map F5 is input into the downsampling module and then stacked three times through the CAFNet Block module to obtain the feature map F6; S7: Pass the feature map F6 through a global average pooling layer, a Layer Norm layer, and a fully connected layer to output a comprehensive classification result of floor color and texture; Among them, the CAFNet Block module includes an asymmetric multi-branch convolution module, a hybrid attention module and a regularization module.
2. The solid wood floor classification and recognition method based on asymmetric multi-branch convolution according to claim 1 is characterized in that: The asymmetric multi-branch convolution module takes the input feature map X m×n×k , along the channel dimension k, it is divided into 3 groups of parallel branches (X m×n×w ,X m×n×h ,X m×n×id ), where X m×n×w is the feature map width branch, X m×n×h is the feature map height branch, X m×n×id is the identity mapping branch, and the number of channels of each branch is g, that is, w = h = id = g; then the sliding window mechanism is used to extract features from each feature map branch, that is, X′ id =X m×n×id In the formula, DWConv is a depth-separable convolution operation, Concat is a concatenation operation, and X' h is the height-direction asymmetric convolution branch output, X' w is the width-wise asymmetric convolution branch output, X' id is the identity mapping branch, 1×k w , k h ×1 is the asymmetric convolution kernel size, and the feature map output by each branch is consistent with the input feature map in height and width; finally, the output connection of each branch is merged, that is, X′=Concat(X′ w ,X′ h ,X′ id ) Where X' is the result of concatenating the outputs of different branches.
3. The solid wood floor classification and recognition method based on asymmetric multi-branch convolution according to claim 1 is characterized in that: The hybrid attention module first increases the dimension of the feature map extracted by the asymmetric multi-branch convolution module through a 1×1 convolution kernel, passes through a GELU activation function, then reduces the dimension through a 1×1 convolution kernel, and then passes through a CBAM attention module, where the CBAM attention module includes a channel attention module and a spatial attention module.
4. The solid wood floor classification and recognition method based on asymmetric multi-branch convolution according to claim 3 is characterized in that: The channel attention module first performs global maximum pooling and global average pooling on the input feature map to obtain two 1×1×C feature vectors, and then sends them to the multi-layer perceptron respectively. The two feature vectors output by the multi-layer perceptron are then added and activated with Sigmoid. The generated feature vector is the attention weight of the channel dimension of the original input feature map. The obtained attention weight of the channel dimension is multiplied by the feature map to be processed to obtain the output feature map of the channel attention module.
5. The solid wood floor classification and recognition method based on asymmetric multi-branch convolution according to claim 3 is characterized in that: The spatial attention module uses the output feature map of the channel attention module as the input feature map. First, the input feature map is subjected to global maximum pooling and global average pooling to obtain two H×W×1 feature maps, which are concatenated in the channel dimension. Then, a convolution kernel of size 7×7 and stride 1 is used for feature extraction, and the dimension is reduced to 1. The spatial attention weight is generated by the Sigmod function, and the feature map is regenerated by multiplying it with the input feature map of the module.
6. The solid wood floor classification and recognition method based on asymmetric multi-branch convolution according to claim 1 is characterized in that: The regularization module is composed of a Layer Scale layer and a Drop Path layer connected in series.
7. The solid wood floor classification and recognition method based on asymmetric multi-branch convolution according to claim 1 is characterized in that: The loss function used for the parameters of each layer and module during training is Focal Loss Pro, and the calculation formula is: FL(p t )=-a t (1-p t ) γ log(p t )-β1log(p HS )-β2log(p HC ) Among them, p t is the model’s predicted probability for the target class, α t is the weight of the sample of the target class, γ is the focus factor, β1 is the weight of the error increased by the patterned dark solid wood floor, β2 is the weight of the error increased by the patterned color difference wood board, p HS is the predicted probability of patterned dark solid wood flooring, p HC is the predicted probability of solid wood flooring with pattern and color difference.
8. The solid wood floor classification and recognition method based on asymmetric multi-branch convolution according to claim 1 is characterized in that: The comprehensive classification results of floor color and texture include patterned white-edged wooden boards, patterned color-difference wooden boards, patterned light-colored wooden boards, patterned dark-colored wooden boards, straight-grained white-edged wooden boards, straight-grained color-difference wooden boards, straight-grained light-colored wooden boards and straight-grained dark-colored wooden boards.
9. The solid wood floor classification and recognition method based on asymmetric multi-branch convolution according to claim 1 is characterized in that: The longitudinal field of view of the collected wooden floor image is at least 1.2 times the width of the wooden floor, the image height pixel number is h, the width pixel number is w, where w is at least 4 times h, and the number of image channels is 3, namely R, G, and B channels.