Feature fusion method based on improved YOLOv8 backbone network
By introducing BiFPN structure and small object detection layer into the YOLOv8 backbone network, the problem of low manual detection efficiency in edge sealing defect detection in wooden boards is solved, and more efficient and accurate defect detection results are achieved.
Patent Information
- Application Number
- CN202510035905.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art relies on manual inspection in the detection of edge sealing defects of wooden boards, resulting in low detection efficiency, high mis-checking and leakage detection rates, which cannot meet the strict requirements of wooden board quality inspection in wooden furniture production lines.
Replace the FPN-PAN structure into the BiFPN structure in the YOLOv8 backbone network, and add a small object detection layer to the detection head. Through cross-scale connection and weighted feature fusion, more shallow features are retained and the performance of small object detection is improved.
Through the combination of BiFPN structure and small object detection layer, the accuracy and efficiency of edge seal defect detection of wooden boards are improved, and the detection ability of defects of different scales is enhanced.
Smart Images

Figure CN120125938A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image feature fusion, and particularly relates to a feature fusion method based on an improved YOLOv8 backbone network. Background Art
[0002] With the development of the times' economy and social progress, residents have an increasing demand for furniture, and the demand for wooden furniture grows steadily every year. Wooden furniture is liked by people because of its advantages such as environmental protection, health, beauty, durability, and high cost performance. The large consumption of wooden furniture brings profits to furniture manufacturers while also posing higher requirements for the product quality of furniture. Wooden boards are the basis for making furniture and a crucial part that affects the quality of wooden furniture. The detection of defects in wooden boards includes multiple steps such as defect identification, defect classification, and defect location. Among them, edge banding is an important part of wooden boards. High-quality edge banding can make the wooden board more beautiful, reduce the erosion of air impurities on the inside of the wooden board and the release of formaldehyde, which is beneficial to environmental protection. During the production process, various factors such as manufacturing processes and wooden board materials can cause edge banding defect problems. The increasingly high requirements for the quality of wooden boards have put forward strict requirements for wooden board manufacturers during the production and quality inspection processes of wooden boards. However, currently, many wooden board production lines still use manual means to detect the edge banding defects and their quality of wooden boards. Because manual detection means have the disadvantages of low detection efficiency, misdetection, and large omission rates, furniture manufacturers need better means to improve the ability to detect the edge banding of wooden boards. Therefore, it is of great significance to study an automated detection means for realizing edge banding defects of wooden boards.
[0003] The neck network mainly performs feature fusion on the multi-scale information obtained by the backbone network. The neck network uses the FPN-PAN structure. The FPN-PAN is a two-way path structure. Through the top-down of PAN, the bottom-up of FPN, and the cross-layer connection path at the same layer, the feature information in the network can be better fused.
[0004] The detection head, as the part of the network that finally outputs the prediction results, adopts a decoupled head structure, separating the class prediction and the location prediction. The Distribution Focal Loss is introduced to calculate the regression loss, and at the same time, the TaskAlignment Learning dynamic matching strategy is used to improve the positive and negative sample matching, enhancing the performance and convergence speed of YOLOv8. Summary of the Invention
[0005] In view of this, the present invention proposes a feature fusion method based on an improved YOLOv8 backbone network, replacing the FPN-PAN structure in the YOLOv8 network with a BiFPN structure and adding a small target detection layer to the detection head.
[0006] Preferably, the YOLOv8 network model is modified to have four detection heads, and an additional upsampling layer Upsample and convolutional operations are added to the neck network.
[0007] Preferably, the BiFPN structure is introduced. In the top-down feature fusion path of the BiFPN structure, nodes with only single-feature information input are removed.
[0008] Preferably, in the bottom-up feature fusion path of the BiFPN structure, cross connections are added between the input nodes and output nodes under the same scale features to fuse feature information while transmitting low-level position information.
[0009] Preferably, the weight fusion formula of the BiFPN structure is:
[0010]
[0011] In the formula, O represents the output, W i , W " represents the weight, e represents the minimum learning rate for constraining numerical oscillation, and I i represents the input.
[0012] Preferably, the output feature map size of the small object detection layer is 160×160, enabling the model to detect small objects with a pixel size of 4×4.
[0013] Compared with the prior art, the feature fusion method based on the improved YOLOv8 backbone network disclosed by the present invention has at least the following beneficial effects:
[0014] 1) In the neck network, the original PAN-FPN network of YOLOv8 is replaced with the BiFPN network. Through cross-scale connection and weighted feature fusion, more shallow features are retained and the parameter calculation amount is reduced.
[0015] 2) The small object detection layer is introduced to extract more small object information and input it into the BiFPN network structure to improve the detection performance of small objects. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] To make the objectives, technical solutions, and beneficial effects of the present invention clearer, the following drawings are provided for illustration:
[0017] Figure 1 It is the network structure diagram of the feature fusion method based on the improved YOLOv8 backbone network in the embodiment of the present invention;
[0018] Figure 2 It is the BiFPN network structure diagram of the feature fusion method based on the improved YOLOv8 backbone network in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0019] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0020] See Figure 1 、 Figure 2 , the network structure diagram of the feature fusion method based on the improved YOLOv8 backbone network of the present invention adds an upsampling Upsample and convolution operation in the neck network of the YOLOv8 network structure, and replaces the FPN-PAN structure with the BiFPN structure. In the top-down feature fusion path of the BiFPN structure, nodes with only single feature information input are deleted. Crosswise connections are added between the input nodes and output nodes under the same scale features to fuse feature information while transmitting low-level position information. A small target detection head is added, and the fused feature information is output to four detection heads.
[0021] The weight fusion formula of the BiFPN structure is:
[0022]
[0023] In the formula, O represents the output, W i , W " represents the weight, e represents the minimum learning rate for restraining numerical oscillation, and I i represents the input.
[0024] YOLOv8 uses the FPN-PAN structure in the Neck layer to achieve feature fusion. FPN has the characteristics of top-down and lateral propagation in information dissemination, and then the PAN structure is used to process the feature information input by FPN. However, the input of PAN only depends on the information output by FPN. Therefore, some of the original feature information lost in the feature extraction process of FPN is not input into the PAN structure, which affects the detection. Compared with the PAN-FPN feature pyramid network, BiFPN adopts cross-scale connection and weighted feature fusion. In the top-down feature fusion path, nodes with only one feature information input are deleted, and only nodes with more feature information input are retained, simplifying the network and enabling better transmission of high-level information to low-level nodes. In the bottom-up feature fusion path, crosswise connections are added between the input nodes and output nodes under the same scale features to fuse more feature information while transmitting low-level position information, so that the fused feature network has higher recognition efficiency. By introducing BiFPN to provide more comprehensive and effective information, the detection ability for defects of different scales is improved.
[0025] In the image of the wooden board edge banding, there are defects of various sizes and shapes. In the original YOLOv8 model, the output feature map sizes are 20×20, 40×40, and 80×80. The feature maps with smaller scales have larger receptive fields and rich semantic information. However, after multiple downsamplings such as deep convolution and pooling in the feature maps, it will also lead to the loss of small target feature information, making the features of local defects not obvious and making it more difficult for the model to detect small targets. On the contrary, the feature maps in the shallow network have larger scales and retain more small target feature information, which is beneficial to the detection of small targets. Aiming at the problems of the YOLOv8 model in the small target detection task, therefore, in the present invention, a small target output layer of 160×160 is introduced on the basis of YOLOv8, so that the model can more accurately detect small targets with a pixel size of 4×4. By directly connecting to the detection head, it is ensured that the target defect information obtained by fusion can be fully utilized to improve the sensitivity of the model for small target detection.
[0026] To verify the improvement effect of the improved module mentioned in the present invention on the model, 4 groups of ablation experiments were designed, and the improved YOLOv8 algorithm was compared with the original YOLOv8 algorithm. The experimental results are shown in Table 1.
[0027] Table 1 Ablation Experiments Based on Improved YOLOv8
[0028]
[0029] From the ablation experiment results in Table 1, taking the mean average precision mAP as an example, the mAP050 value of the YOLOv8n detection algorithm is 60.2%, and the mAP50:95 value is 31.9%. After the BiFPN structure is added to Model 2 alone, the mAP50 value of the model increases by 0.4%, and the mAP50:95 increases by 0.9%. After the small target detection layer is added to Model 3 alone, the mAP50 value of the model is 60.9%, and the mAP50:95 value is 32.5%, which are increased by 0.7% and 0.6% respectively compared with the original algorithm. After Model 4 integrates the BiFPN structure and the small target detection layer, the mAP values of the model reach 62.4% and 33.2%, an increase of 2.2% and 1.3%, showing a relatively large improvement compared with the original model. Through the analysis of the ablation experiment results data, it can be seen the improvement effect of the improved module proposed in the present invention on model detection.
[0030] In addition to the above embodiments, the present invention can also have other implementation manners. All technical solutions formed by equivalent replacement or equivalent transformation are within the protection scope required by the present invention.
[0031] The above has described the present invention in detail, but the specific implementation form of the present invention is not limited thereto. Without departing from the spirit and scope of the claims of this application, those skilled in the art can make various modifications or adaptations.
Claims
1. A feature fusion method based on an improved YOLOv8 backbone network, characterized in that: In the YOLOv8 network, the FPN-PAN structure is replaced with the BiFPN structure, and a small target detection layer is added to the detection head.
2. The feature fusion method based on the improved YOLOv8 backbone network according to claim 1, characterized in that: The YOLOv8 network model is modified to have four detection heads, and an upsample and convolution operation is added to the neck network.
3. The feature fusion method based on the improved YOLOv8 backbone network according to claim 1, characterized in that: The BiFPN structure deletes nodes with only one feature information input in the top-down feature fusion path.
4. The feature fusion method based on the improved YOLOv8 backbone network according to claim 1, characterized in that: The BiFPN structure adds cross-connections between input nodes and output nodes at the same scale feature in the bottom-up feature fusion path, fusing feature information while transmitting low-level position information.
5. The feature fusion method based on the improved YOLOv8 backbone network according to claim 1, characterized in that: The weight fusion formula of the BiFPN structure is: In the formula, O represents output, W i ,W " represents the weight, e represents the minimum learning rate of the constrained numerical oscillation, I i Represents input.
6. The feature fusion method based on the improved YOLOv8 backbone network according to claim 1, characterized in that: The output feature map size of the small object detection layer is 160×160, which enables the model to detect small objects of 4×4 pixel size.