Feature extraction method based on improved YOLOv8 backbone network

By adding convolution and attention fusion module (CAFM) to the YOLOv8 backbone network, the problem of low manual detection efficiency in board edge seal defect detection is solved, more efficient automatic detection capabilities are achieved, and detection accuracy and speed are improved.

CN120088497APending Publication Date: 2025-06-03HANGZHOU DIANZI UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411931579.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art relies on manual means in the detection of edge-sealing defects of wooden boards, resulting in low detection efficiency, high mis-inspection and leakage detection rates, which cannot meet the demand for automated inspection of wooden board production lines.

Method used

The convolution and attention fusion module (CAFM) are added before the fast spatial pyramid pooling module of the YOLOv8 backbone network. This module includes two branches, local and global, and uses the self-attention mechanism to capture global information, and local branches extract local features to improve feature extraction capabilities.

Benefits of technology

By adding CAFM module, the detection capability of the model in the detection of edge-sealing defects in wooden boards is significantly improved, the mAP and regression rate are improved, and the error detection and miss detection rates are reduced, proving the superiority of the CAFM module in this field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088497A_ABST
    Figure CN120088497A_ABST
Patent Text Reader

Abstract

The invention discloses a feature extraction method based on an improved YOLOv8 backbone network, and the method comprises the steps: adding a convolution and attention fusion module in front of a fast spatial pyramid pooling module at the tail end of a backbone network in a YOLOv8 network structure, and enabling the convolution and attention fusion module to comprise a local branch and a global branch, a self-attention mechanism is adopted in the global branch to capture spectral data information, and local features are extracted from the local branch. The convolution and attention fusion module is added to the tail end of the backbone network to enhance the extraction capability of the network for global and local features, the small target detection performance is improved, and feature expression is enhanced to enable the model to automatically learn and reinforce important features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image feature extraction, and particularly relates to a feature extraction method based on an improved YOLOv8 backbone network. Background Art

[0002] YOLOv8 is a YOLO series algorithm based on PyTorch, which was launched in January 2023. It is a currently popular YOLO series detection model and can be used in fields such as image classification and object detection, with excellent detection effects. Compared with the previous generation YOLOv5, YOLOv8 has made many new improvements, improving the performance of the YOLO series in the field of object detection.

[0003] The YOLOv8 network structure consists of three parts: the backbone network BackBone, the neck network Neck, and the detection head Head. The backbone network is used for feature extraction, and its structure consists of a standard convolution module, a C2f module, and an SPPF module. Among them, the standard convolution module is mainly composed of convolution, batch normalization, and the SiLU activation function to improve the generalization ability of the model and accelerate the convergence speed. At the same time, YOLOv8 introduces the C2f module to replace the C3 module used in YOLOv5. The C2f module adds skip connection operations and Split operations compared with the C3 module, and enhances its feature fusion ability in convolution through the residual structure, enabling the gradient of the model to be richer. At the same time, it makes the backbone network more lightweight and improves the inference speed. At the end of the backbone network, the SPPF (Spatial Pyramid Pooling Fast) fast spatial pyramid pooling module is adopted. The SPPF module mainly performs feature fusion through pooling and convolution operations, captures image features at different scales through pooling with different kernel sizes, realizes the superposition and fusion of local features and global features, and adaptively fuses feature information of various scales to enhance the feature extraction ability of the model.

[0004] The neck network is mainly used to perform feature fusion on the multi-scale information obtained by the backbone network. The neck network uses the FPN-PAN structure. The FPN-PAN is a two-way path structure. Through the top-down of PAN, the bottom-up of FPN, and the cross-layer connection path at the same layer, the feature information in the network can be better fused.

[0005] As the part of the network that finally outputs the prediction results, the detection head adopts a decoupled head structure, separating the class prediction and the location prediction. The Distribution Focal Loss is introduced to calculate the regression loss, and at the same time, the TaskAlignment Learning dynamic matching strategy is used to improve the positive and negative sample matching, enhancing the performance and convergence speed of YOLOv8.

[0006] During the production process, various factors such as manufacturing processes and wood board materials can cause edge banding defect problems. The increasingly high quality requirements for wood boards have put forward strict requirements for wood board manufacturers in the production and quality inspection processes of wood boards. However, at present, manual means are still used to detect edge banding defects and quality on many wood-based panel production lines. Because manual inspection methods have the disadvantages of low detection efficiency, misdetection, and large omission rates, furniture manufacturers need better means to improve the ability of wood board edge banding detection. Therefore, the feature extraction of wood board edge banding defects is of great significance in automated detection methods. Summary of the Invention

[0007] In view of this, the present invention proposes a feature extraction method based on an improved YOLOv8 backbone network. A convolution and attention fusion module is added before the fast spatial pyramid pooling module at the end of the backbone network in the YOLOv8 network structure. The convolution and attention fusion module includes local and global branches. In the global branch, a self-attention mechanism is used to capture spectral data information, and the local branch extracts local features.

[0008] Preferably, the local branch feature extraction includes the following steps:

[0009] S10, use 1×1 convolution to adjust the channel dimension;

[0010] S20, perform a channel shuffle operation to further mix and fuse channel information to strengthen information integration;

[0011] S30, the obtained output tensor is concatenated along the channel dimension to generate a new output tensor;

[0012] S40, use 3×3×3 convolution to extract features.

[0013] Preferably, in the channel shuffle operation of S20, the input tensor is divided into several groups along the channel dimension, and depthwise separable convolution is used within each group to trigger channel shuffling.

[0014] Preferably, the process of local branch feature extraction is expressed as:

[0015] F conv =W 3×3×3 (CS(W 1×1 (Y))) (1)

[0016] In the formula, F conv represents the output of the local branch, W 3×3×3 represents 3×3×3 convolution, CS represents the channel shuffle operation, W 1×1 represents 1×1 convolution, and Y represents the input feature.

[0017] Preferably, the global branch feature extraction includes the following steps:

[0018] S11, the input features generate query Q, key K, and value V through 1×1 convolution and three 3×3 depth convolution modules;

[0019] S21, flatten the spatial dimensions of Q, K, and V;

[0020] S31, use the reshaped Q, K, and V to calculate the attention map;

[0021] S41, then add it to the input features through 1×1 convolution.

[0022] Preferably, in the S11, three 3×3 depth convolution modules generate three tensors with the shape of B×H×W×C, where B, H, W, and C respectively represent the number of samples, the height of the feature map, the width of the feature map, and the number of channels.

[0023] Preferably, in the S21, the Softmax function is used to reshape Q, K, and V into the shape of B×N×C, where N = H×W, and B, H, W, and C respectively represent the number of samples, the height of the feature map, the width of the feature map, and the number of channels.

[0024] Preferably, the global branch feature extraction is expressed as:

[0025]

[0026] In the formula, F att represents the output of the global branch, represents the weighted output feature map, represents the weighted value, a represents the learnable scaling parameter, and Y represents the input feature.

[0027] Compared with the prior art, the feature extraction method based on the improved YOLOv8 backbone network disclosed by the present invention has at least the following beneficial effects:

[0028] Add the CAFM attention mechanism to the backbone network to improve the model's ability to capture global and local information. Compared with the effects of TA, CBAM, Simam, and ECA attention in the existing technology on the added baseline model, adding CAFM attention has the greatest improvement on the model's detection ability. Taking the mAP index as a comparison, after adding TA attention, the mAP indexes of the model reached 61.6% and 32.8% respectively, and the regression rate reached 59.2%. Compared with the basic model, the regression rate increased by 1.9%, but both mAP50 and mAP50:95 decreased. After adding CBAM attention, the mAP indexes of the model were 62.4% and 32.8%, and the regression rate reached 58.9%. Compared with the baseline model, mAP50:95 decreased by 0.4%, and the recall rate increased by 1.6%. After adding ECA attention, the recall rate and mAP50 of the model were 57.4% and 62.1% respectively. The recall rate increased by 0.1% compared with the baseline model, and mAP decreased by 0.3%.

[0029] According to the experimental results, compared with the TA, CBAM, and Simam attention mechanisms, after adding CAFM attention, the mAP value and regression rate of the model are in a better position, proving that the CAFM module performs better in the detection of wooden board edge sealing defects. Brief Description of the Drawings

[0030] In order to make the objectives, technical solutions, and beneficial effects of the present invention clearer, the following drawings are provided for the description of the present invention:

[0031] Figure 1 It is the network structure diagram of the feature extraction method based on the improved YOLOv8 backbone network in the embodiment of the present invention;

[0032] Figure 2 It is the convolution and attention fusion module structure diagram of the feature extraction method based on the improved YOLOv8 backbone network in the embodiment of the present invention. Detailed Embodiments

[0033] The preferred embodiments of the present invention will be described in detail below with reference to the drawings.

[0034] See Figure 1 、 Figure 2 , the network structure diagram of the feature extraction method based on the improved YOLOv8 backbone network of the present invention. Due to its local nature and limited receptive field, convolution has deficiencies in modeling global features. In contrast, Transformer with the help of the attention mechanism performs excellently in extracting global features and capturing long-distance dependence relationships. Convolution and attention complement each other in modeling global and local features. The convolution and attention fusion module (CAFM) is a module that combines convolution operations and the attention mechanism. CAFM includes local(Figure 2 Local branch) and global ( Figure 2 Global branch) two branches. In the global branch, the self-attention mechanism is adopted to capture extensive hyperspectral data information, while the local branch focuses on extracting local features to achieve comprehensive denoising. In the YOLOv8 network structure, a convolution and attention fusion module CAFM is added before the SPPF module at the end of the backbone network. This convolution and attention fusion module includes local and global branches. In the global branch, the self-attention mechanism is used to capture spectral data information, and the local branch extracts local features.

[0035] The local branch feature extraction includes the following steps:

[0036] S10, Use 1×1 convolution to adjust the channel dimension;

[0037] S20, Perform channel shuffle operation to further mix and fuse channel information to strengthen information integration. In the channel shuffle operation, the input tensor is divided into several groups along the channel dimension, and depthwise separable convolution is used within each group to trigger channel shuffle;

[0038] S30, The output tensors obtained from each group are concatenated along the channel dimension to generate a new output tensor;

[0039] S40, Use 3×3×3 convolution to extract features.

[0040] The above process of local branch feature extraction is expressed as:

[0041] F conv =W 3×3×3 (CS(W 1×1 (Y))) (1)

[0042] In the formula, F conv represents the output of the local branch, W 3×3×3 represents 3×3×3 convolution, CS represents the channel shuffle operation, W 1×1 represents 1×1 convolution, and Y represents the input feature.

[0043] The global branch feature extraction includes the following steps:

[0044] S11, The input feature generates query Q, key K, and value V through 1×1 convolution and 3 3×3 depth convolution modules, generating 3 tensors with the shape of B×H×W×C, where B, H, W, and C represent the number of samples, feature map height, feature map width, and number of channels respectively;

[0045] S21. To calculate the Attention Map, the spatial dimensions of Q, K, and V need to be flattened, and the Softmax function is used to reshape Q, K, and V into the shape of B×N×C, where N = H×W, and B, H, W, and C represent the number of samples, the height of the feature map, the width of the feature map, and the number of channels, respectively;

[0046] S31. Use the reshaped Q, K, and V to calculate the Attention Map;

[0047] S41. Then, add the result after 1×1 convolution to the input features.

[0048] The above global branch feature extraction is expressed as:

[0049]

[0050] In the formula, F att represents the output of the global branch, represents the weighted output feature map, represents the weighting value, a represents the learnable scaling parameter, and Y represents the input feature.

[0051] Add the CAFM module at the end of the BackBone of the backbone network to enhance the network's ability to extract global and local features, improve the detection performance of small targets, and enhance the feature expression so that the model can automatically learn and strengthen important features.

[0052] In addition to the above embodiments, the present invention may also have other embodiments. Any technical solutions formed by equivalent replacement or equivalent transformation are within the protection scope required by the present invention.

[0053] The above has described the present invention in detail, but the specific implementation form of the present invention is not limited thereto. Without departing from the spirit and scope of the claims of this application, those skilled in the art can make various modifications or adaptations.

Claims

1. A feature extraction method based on an improved YOLOv8 backbone network, characterized in that: In the YOLOv8 network structure, a convolution and attention fusion module is added before the fast spatial pyramid pooling module at the end of the backbone network. The convolution and attention fusion module includes two branches, local and global. The self-attention mechanism is used in the global branch to capture spectral data information, and the local branch extracts local features.

2. The feature extraction method based on the improved YOLOv8 backbone network according to claim 1, characterized in that: The local branch feature extraction comprises the following steps: S10, uses 1×1 convolution to adjust the channel dimension; S20, performing a channel shuffling operation to further mix and fuse channel information to enhance information integration; S30, the obtained output tensors are connected along the channel dimension to generate a new output tensor; S40, uses 3×3×3 convolution to extract features.

3. The feature extraction method based on the improved YOLOv8 backbone network according to claim 2, characterized in that: In the channel shuffling operation of S20, the input tensor is divided into several groups along the channel dimension, and a depthwise separable convolution is used in each group to induce channel shuffling.

4. The feature extraction method based on the improved YOLOv8 backbone network according to claim 2, characterized in that: The process of extracting local branch features is expressed as: F conv =W 3×3×3 (CS(W 1×1 (Y))) (1) In the formula, F conv represents the output of the local branch, W 3×3×3 represents a 3×3×3 convolution, CS represents a channel shuffle operation, and W 1×1 represents 1×1 convolution, and Y represents the input feature.

5. The feature extraction method based on the improved YOLOv8 backbone network according to claim 1, characterized in that: The global branch feature extraction comprises the following steps: S11, the input features are passed through 1×1 convolution and three 3×3 deep convolution modules to generate query Q, key K, and value V; S21, flatten the spatial dimensions of Q, K, and V; S31, use the reshaped Q, K, V to calculate the attention map; S41, then adds the input features through 1×1 convolution.

6. The feature extraction method based on the improved YOLOv8 backbone network according to claim 5, characterized in that: In the S11, three 3×3 depth convolution modules are used to generate three tensors of shape B×H×W×C, where B, H, W, and C represent the number of samples, feature map height, feature map width, and number of channels, respectively.

7. The feature extraction method based on the improved YOLOv8 backbone network according to claim 6, characterized in that: In S21, the Softmax function is used to reshape Q, K, and V into a shape of B×N×C, where N=H×W, and B, H, W, and C represent the number of samples, feature map height, feature map width, and number of channels, respectively.

8. The feature extraction method based on the improved YOLOv8 backbone network according to claim 7, characterized in that: The global branch feature extraction is expressed as: In the formula, F att represents the output of the global branch, represents the weighted output feature map, represents the weight value, a represents the learnable scaling parameter, and Y represents the input feature.