Wood repairing method and device and medium

By improving the backbone network, neck network and detection head design of the YOLOv8 model, combined with deep separation convolution and lightweight attention module, the problem of inefficiency in wood defect detection and repair is solved, high-precision automated repair is achieved, adapting to complex production environments, and improving repair effect and production efficiency.

CN120411054AActive Publication Date: 2025-08-01KUNMING UNIV OF SCI & TECH
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510553714.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-01
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The existing wood defect detection and repair technology is inefficient, relies on manual operations and cannot achieve high-precision automated repair. The traditional YOLO model lacks specific optimization for wood defects, resulting in unstable detection results and unsatisfactory repair results.

Method used

The improved YOLOv8 model is adopted, and the Conv module of the backbone network is replaced by the MASFAConv module, the C2f module of the neck network is C2f_DSCA module, and a P2 detection head is added to the detection head, combining weighted box fusion and multi-scale detection fusion strategies to improve defect detection accuracy; using deep separable convolution and lightweight attention modules to enhance the ability to capture wood surface texture and small defects, and combining the ModbusTCP protocol and the PLC control system to achieve automated repair.

Benefits of technology

It improves the accuracy of wood defect detection and the efficiency of automated repair, can achieve efficient and accurate defect repair in complex production environments, enhances the detection ability of tiny defects on the surface of wood, and ensures consistency of repair quality and production efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411054A_ABST
    Figure CN120411054A_ABST
Patent Text Reader

Abstract

The invention relates to a wood repairing method and device and a medium, and belongs to the technical field of wood repairing. The method comprises the following steps: acquiring a real-time wood image; the real-time wood image is input into a large wood defect detection model, a defect bounding box and a defect type are obtained, and the large wood defect detection model is based on an improved YOLOv8 model; replacing a Conv module in a backbone network of the initial YOLOv8 model with a MASFAConv module, replacing a C2f module in a neck network with a C2fDSCA module, and adding a P2 detection head in a detection head to obtain an improved YOLOv8 model; performing weighted average processing on the defect bounding box by adopting weighted box fusion and multi-scale detection fusion strategies to obtain a target defect bounding box, and generating position information of the target defect bounding box at the same time; and controlling a repairing device to repair the wood according to the position information of the target defect bounding box and the defect type. According to the method, wood defects can be accurately positioned, and the wood repairing quality is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the technical field of wood repair, and particularly relates to a wood repair method, device and medium. Background Art

[0002] As a commonly used building and industrial material, the surface defects of wood (such as cracks, dead knots, live knots, etc.) will seriously affect the quality and service life of products. Traditional wood defect detection methods mainly rely on manual visual inspection and simple machine vision technology. However, manual detection is not only inefficient, but also easily affected by human factors, resulting in missed or misjudged detections. Currently, wood defect repair usually relies on manual operation or semi-automated repair techniques, and mostly uses repair materials such as putty.

[0003] Traditional repair methods rely on manual operation to judge the specific situation of defects and manually select putty for repair, which is not only inefficient, but also prone to uneven repair effects and unable to ensure the consistency of repair quality. In addition, the existing repair techniques cannot achieve large-scale automated applications in a high-efficiency production environment, and there are many challenges in the repair process.

[0004] Currently, computer vision and image processing methods have been gradually introduced into wood defect detection technology. Many traditional methods rely on techniques such as edge detection and morphological processing. These methods lack sufficient robustness and adaptability to the surface texture, light changes and complexity of wood defects, resulting in unstable detection results and being difficult to meet the high-precision requirements to achieve full coverage of the repair path.

[0005] In recent years, deep learning, especially the YOLO series of algorithms, has been applied to wood defect detection. Although the YOLO algorithm performs well in real-time detection and accuracy, most of the existing YOLO models are for general object detection tasks and lack specific optimization for wood defects. Therefore, its application in wood defect detection still faces challenges.

[0006] In order to improve the efficiency and quality of wood repair, automated repair technology has gradually become a research hotspot. Some existing automatic repair systems attempt to achieve the automation of the repair process through robotics technology. Although it has been applied in some simple repair tasks, most systems still lack precise defect detection and positioning capabilities, resulting in unsatisfactory repair effects. Summary of the Invention

[0007] The present disclosure proposes a wood repair method, device and medium to solve the above technical problems.

[0008] According to a first aspect of the present disclosure, a wood repair method is provided, the method comprising: acquiring a real-time image of wood; inputting the real-time image of wood into a large wood defect detection model to obtain a defect bounding box and a defect type, wherein the large wood defect detection model is based on an improved YOLOv8 model; replacing the Conv module in the backbone network of the initial YOLOv8 model with a MASFAConv module, replacing the C2f module in the neck network with a C2f_DSCA module, and adding a P2 detection head in the detection head to obtain the improved YOLOv8 model; performing weighted average processing on the defect bounding box by using a weighted box fusion and multi-scale detection fusion strategy to obtain a target defect bounding box, and simultaneously generating target defect bounding box position information; and controlling a repair device to repair the wood according to the target defect bounding box position information and the defect type.

[0009] In some embodiments, the processing of the real-time image of wood by the large wood defect detection model includes: expanding the real-time image of wood by using depthwise separable convolution to obtain two initial branch features; processing the initial features in the branches by using a sub-module composed of a series connection of bottleneck structure modules of multiple depthwise separable convolutions, each module internally using two layers of depthwise separable convolution, including a depthwise separable convolution for feature compression and a depthwise separable convolution for feature expansion, and when the number of input channels is the same as the number of output channels, using a residual connection, and after multiple series processing, forming multiple local branch features; splicing the two initial branch features and the multiple local branch features in the channel dimension; performing adaptive reweighting on the spliced features by using a convolution-based lightweight attention module, the module using two layers of convolution, the first layer of convolution is dimension-reduced and activated by SiLU, the second layer of convolution is dimension-increased and then passed through a Sigmoid activation function to obtain channel attention weights, and finally multiplying the weights with the original spliced features channel by channel; and mapping the features weighted by the adaptive weights to the target output channels by using one layer of depthwise separable convolution.

[0010] In some embodiments, before expanding the real-time image of wood by using one layer of depthwise separable convolution, it further includes: performing global average pooling and max pooling on the real-time image of wood by using channel attention, and generating a channel mask through convolution layer mapping; performing global average pooling and max pooling on the real-time image of wood by using spatial attention, and generating a spatial mask through convolution mapping; fusing the channel mask and the spatial mask by using a weighting coefficient α to obtain a preliminarily reweighted output; inputting the preliminarily reweighted output into the channel attention and spatial attention again to obtain a new channel mask and a new spatial mask; and fusing the new channel mask and the new spatial mask by using a weighting coefficient β to obtain a finally reweighted output.

[0011] In some embodiments, the defect types include cracks and knots.

[0012] According to a second aspect of the present disclosure, there is provided a wood repair device, comprising: a real-time wood image acquisition module for acquiring a real-time wood image; a wood defect identification module for inputting the real-time wood image into a large wood defect detection model to obtain a defect bounding box and a defect type, wherein the large wood defect detection model is based on an improved YOLOv8 model; replacing the Conv module in the backbone network of the initial YOLOv8 model with a MASFAConv module, replacing the C2f module in the neck network with a C2f_DSCA module, and adding a P2 detection head to the detection head to obtain the improved YOLOv8 model; a target defect bounding box position information generation module for performing weighted average processing on the defect bounding box by adopting a weighted box fusion and multi-scale detection fusion strategy to obtain a target defect bounding box, and simultaneously generating target defect bounding box position information; and a wood repair control module for controlling the repair device to repair the wood according to the target defect bounding box position information and the defect type.

[0013] According to a third aspect of the present disclosure, there is provided a wood repair device, comprising: a memory; and a processor coupled to the memory, the processor being configured to execute the wood repair method as described above based on instructions stored in the memory.

[0014] According to a fourth aspect of the present disclosure, there is provided a computer-readable storage medium having computer program instructions stored thereon, and when the instructions are executed by a processor, the wood repair method as described above is implemented.

[0015] By adopting the above technical solution, the embodiments of the present disclosure can achieve the following beneficial technical effects: The present disclosure is based on the YOLOv8 model framework, and has made targeted improvements in the design of the model backbone, neck network and detection head, fully considering the actual situation of complex wood surface texture and diverse defect sizes. In the backbone network, a new MASFConv module is used to replace the original Conv module. In the neck network, part of the original C2f module is replaced by the improved C2f_DSCA module. While maintaining the advantages of multi-branch feature fusion, the improved C2f_DSCA module adopts the design of depthwise separable convolution, namely DSC, and changes all the original 1×1 convolutions to 3×3 convolutions to improve the perception field, thereby expanding the perception field and enhancing the ability to capture local spatial information. In order to further improve the model's sensitivity to local defects (such as long and thin cracks and wormholes), a lightweight attention module based on convolution is introduced inside the C2f_DSCA module. This module utilizes two layers of 3×3 convolutions, employing SiLU and Sigmoid activation functions, to generate channel-adaptive attention weights. This reweights the concatenated features from multiple branches, thereby highlighting the characteristics of defective areas. To address the shortcomings of the traditional YOLOv8 model in small object detection, this solution incorporates a new P2 detection head. The P2 detection head is primarily responsible for capturing low-level, fine-grained feature information. By fusing the P2 layer output with other high-level features, it improves the accuracy of detecting small defects on wood surfaces.

[0016] Depthwise Separable Convolution (DSC) employs a 3×3 convolution kernel to perform depthwise and pointwise convolution. Compared to traditional convolution, DSC significantly reduces the number of parameters and computational complexity while maintaining strong feature extraction capabilities, effectively capturing local texture information in wood cracks and insect holes. Furthermore, DSC utilizes a bottleneck structure to achieve feature compression and restoration through two layers of 3×3 depthwise separable convolutions, employing residual connections when the input and output channels are identical. This module, stacked repeatedly in a multi-branch structure, deeply mines local details, helping to enhance the representation of subtle defects. Finally, two layers of 3×3 convolutions, activated with SiLU and Sigmoid activation functions, achieve channel-adaptive reweighting. This allows for reweighting of the concatenated multi-branch features, highlighting key defect information and suppressing irrelevant background noise. In the post-processing stage, this disclosure incorporates the weighted box fusion (WBF) algorithm to weightedly fuse the detection boxes output by each network layer, reducing duplicate boxes and improving detection accuracy. It also supports multi-scale input for inference on the same image, ultimately fusing the results to accommodate production scenarios with complex wood surface patterns and large variations in lighting. Regarding automated repair, this paper combines the Modbus TCP protocol with YOLOv8 defect detection path coordinates and a PLC-controlled repair system, enabling efficient and accurate transmission of defect data and issuance of repair instructions. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings forming a part of the specification depict embodiments of the present disclosure and, together with the specification, are used to explain the principles of the present disclosure.

[0018] With reference to the drawings, the present disclosure can be more clearly understood according to the following detailed description.

[0019] Figure 1 is one of the flowcharts showing a wood repair method according to some embodiments of the present disclosure.

[0020] Figure 2 is the second flowchart showing a wood repair method according to some embodiments of the present disclosure.

[0021] Figure 3 is a diagram showing the overall architecture of the YOLOv8 model according to some embodiments of the present disclosure.

[0022] Figure 4 is a diagram showing the overall architecture of the C2f_DSCA module according to some embodiments of the present disclosure.

[0023] Figure 5 is a schematic diagram showing a residual connection according to some embodiments of the present disclosure.

[0024] Figure 6 is a diagram showing the architecture of MSWFAConv according to some embodiments of the present disclosure.

[0025] Figure 7 is a diagram showing the architecture of MSWFA according to some embodiments of the present disclosure.

[0026] Figure 8 is a block diagram showing a wood repair device according to some embodiments of the present disclosure.

[0027] Figure 9 is a block diagram showing a wood repair device according to some other embodiments of the present disclosure.

[0028] Figure 10 is a block diagram showing a computer system for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] Now, various exemplary embodiments of the present disclosure will be described in detail with reference to the drawings. It should be noted that: Unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.

[0030] Meanwhile, it should be understood that, for the convenience of description, the dimensions of the various parts shown in the drawings are not drawn in actual proportional relationship. The following description of at least one exemplary embodiment is actually only illustrative and in no way restricts the present disclosure and its application or use.

[0031] Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the said technologies, methods, and devices should be regarded as part of the specification.

[0032] In all the examples shown and discussed here, any specific values should be construed as merely exemplary and not as a limitation. Therefore, other examples of the exemplary embodiments may have different values.

[0033] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0034] Figure 1 is one of the flowcharts showing a wood repair method according to some embodiments of the present disclosure. As Figure 1 shown, the wood repair method includes steps S110 to S140.

[0035] In step S110, a real-time image of the wood is acquired.

[0036] In step S120, the real-time image of the wood is input into a large wood defect detection model to obtain a defect bounding box and a defect type. Among them, the wood defect detection model is based on the improved YOLOv8 model; the Conv module in the backbone network of the initial YOLOv8 model is replaced with the MASFAConv module, the C2f module in the neck network is replaced with the C2f_DSCA module, and a P2 detection head is added to the detection head to obtain the improved YOLOv8 model. Global average pooling and max pooling are performed on the real-time image of the wood using channel attention, and a channel mask is generated through convolution layer mapping;

[0037] The processing of the large wood defect detection model for real-time wood images includes: performing global average pooling and max pooling on the real-time wood image using spatial attention, and generating a spatial mask through convolutional mapping; fusing the channel mask and the spatial mask with a weighting coefficient α to obtain the output of the first re-weighting; inputting the output of the first re-weighting into the channel attention and spatial attention again to obtain a new channel mask and a new spatial mask; fusing the new channel mask and the new spatial mask with a weighting coefficient β to obtain the output of the final re-weighting; expanding the real-time wood image using depthwise separable convolution to obtain two initial branch features; processing the initial features in the branch using a sub-module composed of a series of bottleneck structure modules of multiple depthwise separable convolutions. Each module internally uses two layers of depthwise separable convolution, including depthwise separable convolution for feature compression and depthwise separable convolution for feature expansion. When the number of input channels is the same as the number of output channels, a residual connection is used. After multiple series of processing, multiple local branch features are formed; splicing the two initial branch features and multiple local branch features in the channel dimension; using a lightweight attention module based on convolution to perform adaptive re-weighting on the spliced features. This module uses two layers of convolution. After the first layer of convolution reduces the dimension and is activated by SiLU, the second layer of convolution increases the dimension and obtains the channel attention weight through the Sigmoid activation function. Finally, the weight is multiplied by the original spliced features channel by channel; mapping the features weighted by the adaptive weight to the target output channels through one layer of depthwise separable convolution.

[0038] This disclosure is based on the YOLOv8 model framework and has made targeted improvements and designs in the model backbone, neck network, and detection head design, fully considering the actual situation of complex wood surface textures and diverse defect sizes.

[0039] As Figure 3 shown, in the backbone network, a new MASFConv module is used to replace the original Conv module. In the neck network, some of the original C2f modules are replaced by the improved C2f_DSCA module. As Figure 4 shown, while maintaining the advantages of multi-branch feature fusion, the improved C2f_DSCA module adopts the design of depthwise separable convolution, i.e., DSC, and all the original 1×1 convolutions used to increase the receptive field are changed to 3×3 convolutions, thereby expanding the receptive field and enhancing the ability to capture local spatial information. To further improve the sensitivity of the model to local defects (such as slender cracks and insect holes), a lightweight attention module based on convolution is introduced inside the C2f_DSCA module. This module uses two layers of 3×3 convolution, adopts the SiLU activation function and the Sigmoid activation function to generate channel adaptive attention weights, and re-weights the features after multi-branch splicing, thus highlighting the features of the defect area.

[0040] To address the shortcomings of the traditional YOLOv8 model in small object detection, this solution adds a P2 detection head. The P2 detection head is primarily responsible for capturing low-level, fine-grained feature information. By fusing the P2 layer output with other high-level features, it improves the accuracy of detecting small defects on wood surfaces.

[0041] Depthwise separable convolution uses a 3×3 convolution kernel to perform depthwise convolution and pointwise convolution. Compared with traditional convolution, DSC significantly reduces the number of parameters and computational complexity while maintaining strong feature extraction capabilities, effectively capturing local texture information in wood cracks and insect holes. Furthermore, depthwise separable convolution is used to build a bottleneck structure, achieving feature compression and recovery through two layers of 3×3 depthwise separable convolution, and using residual connections when the input and output channels are the same, such as Figure 5 As shown in the figure, this module is repeatedly stacked in a multi-branch structure, deeply exploring local details and helping to enhance the feature representation of small defects. Finally, two layers of 3×3 convolutions, activated by SiLU and Sigmoid activation functions, achieve channel-adaptive reweighting. This allows the concatenated multi-branch features to be reweighted, highlighting key defect information and suppressing irrelevant background noise.

[0042] In step S130 , a weighted average process is performed on the defect bounding box using a weighted box fusion and multi-scale detection fusion strategy to obtain a target defect bounding box and generate target defect bounding box position information.

[0043] This paper combines the weighted box fusion (WBF) algorithm to perform weighted fusion on the detection boxes output by each layer of the network, reducing duplicate boxes and improving detection accuracy. It also supports multi-scale input (for example, 640, 960, 1280, etc.) to reason about the same image and finally fuse the results to adapt to production scenarios with complex wood surface patterns and large changes in lighting.

[0044] In step S140 , a repairing device is controlled to repair the wood according to the target defect boundary box position information and the defect type.

[0045] In terms of automated repair, this paper combines the Modbus TCP protocol with the YOLOv8 defect detection path coordinates and the PLC control repair system to achieve efficient and accurate defect data transmission and repair instruction issuance.

[0046] like Figure 2As shown, the system connects to the PLC via the Modbus TCP protocol using the IP address. After the program starts, an independent thread reads the status of PLC coil 0 at regular intervals through Modbus communication. When the read coil address changes from 0 to 1, the program sets the boolean value for triggering the photo taking to True, indicating that a photo taking detection process is required. The current frame image will be obtained from the Hikvision industrial camera. After the acquisition is completed, the boolean value for triggering the photo taking is reset to False. While obtaining the current frame image, the position of the current PLC on the production line is read, and then the detection thread is started to perform defect detection processing on the obtained current frame image. Inside the detection thread, the YOLO model is used to detect the obtained image to obtain the defect bounding box coordinates and category information. For each detected defect, the conversion from camera pixels to the actual physical coordinates of the field of view is performed based on the four corner coordinates of the defect. Then, based on the actual physical coordinates of the defect, a trajectory planning algorithm is used to ensure the optimal glue output and coverage area each time the valve is opened, and the defect area can be effectively covered within a limited time, significantly enhancing the repair effect and production efficiency.

[0047] The information such as the coordinates of the wood defect positions detected by the YOLOv8 model is transmitted to the PLC in real time in a standardized format, that is, in the form of 16-bit integers, so as to control the PLC to perform the repair task. The specific glue filling interval will be set by changing the step size. After all the coordinates are transmitted, the position of the production line read is sent to a specific address of the PLC. After all the data is sent, a completion signal is sent to a specific completion signal address to achieve a closed-loop of the entire set of processes. At the same time, the system has a friendly human-machine interface, providing real-time display of detection results and system status monitoring. For possible abnormal situations, such as communication failures, camera errors, etc., the system designs a perfect abnormal handling mechanism to ensure the stable operation of the system.

[0048] For the wood defect detection scenario, the overall architecture of the YOLOv8 model was improved and optimized. First, a large number of wood surface defect images were collected at the industrial site, and the defects in them were labeled using annotation tools. The defect types were divided into two categories: cracks and knots. The dataset was expanded through data augmentation methods such as rotation, cropping, and brightness adjustment to improve the generalization ability of the model. Then, the hyperparameter optimization tool Optuna was used to automatically search for the optimal training parameters such as the learning rate, batch size, and number of training epochs to obtain the YOLOv8 model with the best performance. On this basis, aiming at the problems of complex wood surface texture and difficult detection of small defect targets, multiple improvements were made to the network structure of YOLOv8. First, in the backbone network, MASFConv was used to replace Conv to further improve the detail capture ability and defect localization accuracy in the wood board defect recognition process. The designed C2f_DSCA module was used to replace some C2f modules in the neck network of the original model to enhance the feature extraction ability for small defects. The designed C2f_DSCA module is a structure that combines depthwise separable convolution and convolutional attention mechanism, taking into account both lightweight and feature expression capabilities. It makes the model more sensitive to small targets and has richer feature expressions without significantly increasing the number of parameters through multi-branch feature fusion and channel attention reweighting, thus effectively improving the model's response ability to small defects in complex backgrounds. At the same time, a lightweight attention mechanism was introduced to further emphasize the features of the defect area. While fusing multi-scale information, this attention module only brings minimal computational overhead but significantly improves the feature expression ability, enabling the model to focus more on the wood defect area without being disturbed by background noise. Finally, a new P2 branch was added, and finally there are four detection branches, which can better cover various scales from small targets to large targets.

[0049] In some embodiments, mainly aiming at the C2f module in the backbone of the YOLOv8 model, a newly designed C2f_DSCA module was proposed. This module uses a full-process 3×3 (3 rows by 3 columns in height and width of the convolution kernel) depthwise separable convolution (DSC) and independently designs and integrates a lightweight attention mechanism based on convolution (ConvAtt) to enhance the feature capture ability for small target defects such as cracks and wormholes on the wood surface, while reducing the number of model parameters and computational complexity.

[0050] The first step, feature expansion and branch initialization: On the left branch of the module, first use a 3×3 depthwise separable convolution (DSC) to expand the input features. This module consists of two parts. The first part is the depth convolution, that is, perform 3×3 convolution on each input channel respectively to extract local spatial information; then use another 3×3 convolution, that is, change the original pointwise convolution to 3×3 convolution to achieve channel transformation.

[0051] This step maps the input features to twice the hidden channels, and then divides them equally according to the channel dimension to obtain two initial feature branches, laying the foundation for subsequent multi-branch fusion.

[0052] The second step involves local feature extraction and residual fusion: In the right branch, a submodule consisting of a bottleneck module consisting of n depthwise separable convolutions connected in series is used to further process a subset of features within the branch. Each module employs two layers of 3×3 depthwise separable convolutions: the first layer implements feature compression, and the second layer implements feature expansion. When the number of input and output channels is the same, a residual connection is also implemented within the module, directly adding the input to the output to mitigate information loss. After n series processing, multiple local feature branches are formed, which can deeply explore the subtle texture information of wood defect areas. n can be modified and adapted based on computer performance and computational requirements.

[0053] The third step is to combine multi-branch feature splicing and convolutional attention fusion: The two parts of the initial features obtained in the first step are combined with the branch features obtained in the second step through the bottleneck structure module of the depthwise separable convolution, and then spliced in the channel dimension. The total number of channels of the spliced features is (2 + n) times the number of hidden channels. To further enhance the important information of the defect area, a lightweight convolution-based attention module is designed to adaptively reweight the spliced features. This module uses two layers of 3×3 convolution. The first layer of convolution is activated by SiLU after dimensionality reduction. The second layer of convolution is activated by Sigmoid activation function after dimensionality increase to obtain channel attention weights. Finally, the weights are multiplied by the original spliced features channel by channel to achieve dynamic adjustment of information.

[0054] Step 4: Output Mapping and Feature Integration: In the right branch, the features reweighted by the convolution-based lightweight attention module are mapped to the target output channels through a layer of 3×3 depthwise separable convolution. This layer not only adjusts the number of channels but also further integrates the feature information from each branch, ensuring that the output feature map contains rich local details and has high discriminative power. The features output by this module serve as input to the subsequent detection head for the final prediction of defect boxes.

[0055] In some embodiments, as Figure 6 and Figure 7As shown, in order to further improve the detail capture ability and defect localization accuracy in the process of wooden board defect recognition, multi-stage attention is added on the basis of the conventional convolutional neural network, and a weighted fusion method is used to balance the contributions of channel attention and spatial attention. Through continuous iteration and optimization, in the recognition of wooden board defects, it has a significant improvement in the detection effect of subtle features such as knots (live knots and dead knots). A module named MASFAConv (Multi Stage Attention Fusion Attention Convolution) is designed to replace the Conv module in the backbone.

[0056] The advantages are that it avoids introducing too many parameters and computational amounts, ensures smooth operation on a conventional hardware platform, and retains the lightweight characteristics. It combines the information aggregation at the channel level with the attention mechanism at the spatial level, and iterates multiple times to avoid information omission. Through the adjustable weight coefficients α and β, the influence ratio of the two-stage attention on feature extraction can be flexibly adjusted.

[0057] The first step, multi-stage attention and weighted fusion: Before entering the multi-branch feature splicing and global attention reweighting, in this embodiment, the features are initially processed and multiplicatively weighted through two major steps, namely multi-stage attention branch initialization and layer-by-layer attention enhancement and weighted fusion.

[0058] The first stage: In the forward propagation function of the model, the input feature x is first extracted to perform channel attention and spatial attention. The channel attention obtains statistics by performing global average pooling and max pooling on the input feature respectively, and then generates a channel mask through mapping by a fully connected or convolutional layer; the spatial attention obtains a spatial mask through convolution after taking the maximum and average of the feature map along the channel dimension. Then, through the defined weighting coefficient a, fusion is performed between the channel and spatial attention to obtain the output x1 of the first reweighting. At this time, the network can initially highlight the channel and spatial position information corresponding to the potential defect areas.

[0059] The second stage: The output x1 of the first step is input into the channel attention and spatial attention modules again, and new attention weights are similarly extracted; through the defined weighting coefficient β, further balance is achieved between the channel and spatial attention results, and they are multiplied element-wise with x1 to form a new output x2; this process is equivalent to strengthening the network's perception of key textures and regions to some extent, making it easier to capture the subtle features of knots. In this stage and subsequent branch operations, residual connections will be adopted, that is, when the number of channels matches, the input is directly superimposed on the output, so as to ensure smooth information flow between the front and back features and alleviate the problem of gradient disappearance or detail loss that may occur in the deep network.

[0060] Under the continuous action of the above two stages, the model will continuously stack the attention effects from the lower levels to the higher levels, and flexibly allocate channel and spatial information through the adjustment of α and β. The output features of the first step will be concatenated in multiple branches together with the features from other branches, and then further reweighted through the global attention module to obtain better defect detection performance.

[0061] Second step, multi-branch feature aggregation: In the two stages of the first step, after multi-stage attention enhancement and weighting processing, multiple feature branches with different representation capabilities are formed. Stack all the branch outputs according to the channel dimension. Let the initial number of branches be k, and the number of channels of each branch output be C f , then the total number of channels of the concatenated feature map is k*C f , and it remains the same as the original branch in the spatial size H*W. This concatenation operation can effectively integrate information from different attention perspectives and provide a richer feature basis for subsequent global attention.

[0062] Third step, dimensionality reduction convolution and non-linear mapping: To alleviate the problem of excessive number of feature channels and strengthen the expression of key features during fusion, a 3×3 convolution operation is first used to reduce the dimension of the concatenated large-channel number features. Let the dimensionality reduction factor be r, and the number of channels after dimensionality reduction is k*C f / r. At the same time, the SiLU non-linear activation function is used to add representation ability to make the subsequent attention generation process more flexible. This process is equivalent to extracting a more discriminative compressed representation on the high-dimensional concatenated features to reduce the computational amount and reduce the interference of redundant features.

[0063] Fourth step, dimensionality increase convolution and channel attention inference: After obtaining the intermediate features after dimensionality reduction, another convolution is used for dimensionality increase to restore the number of channels to the original scale, and the Sigmoid activation function is combined to generate the channel attention weight map. To balance lightweight and sufficient expression, the Sigmoid activation is selected to construct the channel mask. Specifically, the size of the convolution output is the same as the concatenated features, and the importance of the current channel is represented by the per-channel weight coefficient. This process is called the channel attention mechanism, which can significantly highlight the channel features contributing to the target defect area and suppress the noise channels.

[0064] Fifth step, global feature reweighting and output: Multiply the channel attention weights obtained in the fourth step with the concatenated features channel by channel to achieve global redistribution of the fused feature map. This multiplication operation can be regarded as adaptively amplifying or weakening different channels to make the network focus on identifying and locating the key patterns of the defect area. At the same time, to ensure that the final output can be closely connected to the subsequent tasks, a 3×3 convolution will be inserted additionally after this process to perform a matching mapping on the number of channels. Let k*C fMapping back to the channel dimension required by the backbone network or detection head. At this point, the multi-branch features can be reweighted by global attention to form a unified, more discriminative high-dimensional representation for subsequent defect detection or classification.

[0065] Figure 8 is a block diagram illustrating a wood repairing device according to some embodiments of the present disclosure.

[0066] like Figure 8 As shown, the wood repairing device 800 includes a wood real-time image acquisition module 810 , a wood defect recognition module 820 , a target defect bounding box position information generation module 830 , and a wood repairing control module 840 .

[0067] The wood real-time image acquisition module 810 is configured to acquire a real-time image of the wood;

[0068] The wood defect recognition module 820 is configured to input the real-time wood image into a large wood defect detection model to obtain a defect bounding box and defect type, wherein the large wood defect detection model is based on an improved YOLOv8 model; the improved YOLOv8 model is obtained by replacing the Conv module in the backbone network of the initial YOLOv8 model with a MASFAConv module, replacing the C2f module in the neck network with a C2f_DSCA module, and adding a P2 detection head to the detection head;

[0069] The target defect bounding box position information generating module 830 is configured to perform weighted averaging processing on the defect bounding box using weighted box fusion and multi-scale detection fusion strategies to obtain the target defect bounding box and generate the target defect bounding box position information at the same time;

[0070] The wood repair control module 840 is configured to control the repair device to repair the wood according to the target defect boundary box position information and the defect type.

[0071] Figure 9 FIG. 1 is a block diagram showing a wood repairing device according to other embodiments of the present disclosure. Figure 9 As shown, wood repairing device 900 includes a memory 910 and a processor 920 coupled to memory 910. Memory 910 is configured to store instructions for executing corresponding embodiments of a wood repairing method. Processor 920 is configured to execute the wood repairing method according to any of the embodiments of the present disclosure based on the instructions stored in memory 910.

[0072] Figure 10 is a block diagram illustrating a computer system for implementing some embodiments of the present disclosure.

[0073] like Figure 10As shown, the computer system 1000 may be embodied in the form of a general-purpose computing device. The computer system 1000 includes a memory 1010, a processor 1020, and a bus 1030 that connects different system components.

[0074] The memory 1010 may include, for example, a system memory, a non-volatile storage medium, etc. The system memory stores, for example, an operating system, application programs, a boot loader, and other programs. The system memory may include a volatile storage medium, such as random access memory (RAM) and / or cache memory. The non-volatile storage medium stores, for example, instructions for executing corresponding embodiments of at least one of the wood repair methods. The non-volatile storage medium includes, but is not limited to, disk memory, optical memory, flash memory, etc.

[0075] The processor 1020 may be implemented in the form of a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates, or discrete hardware components such as transistors. Accordingly, each module such as a wood real-time image acquisition module, a wood defect identification module, a target defect bounding box position information generation module, and a wood repair control module may be implemented by a central processing unit (CPU) running instructions for executing corresponding steps in the memory, or may be implemented by dedicated circuits for executing the corresponding steps.

[0076] The bus 1030 may use any of a variety of bus structures. For example, the bus structure includes, but is not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus.

[0077] The computer system 1000 may further include an input / output interface 1040, a network interface 1050, a storage interface 1060, etc.

[0078] These interfaces 1040, 1050, 1060, as well as the memory 1010 and the processor 1020, may be connected via the bus 1030. The input / output interface 1040 provides a connection interface for input / output devices such as a display, a mouse, and a keyboard. The network interface 1050 provides a connection interface for various networking devices. The storage interface 1060 provides a connection interface for external storage devices such as a floppy disk, a USB flash drive, and an SD card.

[0079] Here, various aspects of the present disclosure have been described with reference to the flowcharts and / or block diagrams of methods, apparatuses, and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of the blocks, may be implemented by computer-readable program instructions.

[0080] These computer-readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable apparatus to produce a machine, such that the instructions executed by the processor create means for implementing the functions specified in one or more boxes of the flowchart and / or block diagram.

[0081] These computer-readable program instructions may also be stored in a computer-readable memory, such instructions causing a computer to operate in a particular manner, thereby producing a manufacture including instructions for implementing the functions specified in one or more boxes of the flowchart and / or block diagram.

[0082] The present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects.

[0083] So far, the wood repair method, apparatus, and medium according to the present disclosure have been described in detail. To avoid obscuring the concept of the present disclosure, some details known in the art have not been described. Those skilled in the art can fully understand how to implement the technical solutions disclosed herein based on the above description.

[0084] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art should understand that the above examples are for illustrative purposes only and not for limiting the scope of the present disclosure. Those skilled in the art should understand that the above embodiments may be modified without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.

Claims

1. A wood repair method, characterized in that, The method includes: Obtain the real-time image of the wood; Input the real-time image of the wood into the large wood defect detection model to obtain the defect bounding box and the defect type. Among them, the large wood defect detection model is based on the improved YOLOv8 model. Replace the Conv module in the backbone network of the initial YOLOv8 model with the MASFAConv module, replace the C2f module in the neck network with the C2f_DSCA module, and add the P2 detection head in the detection head to obtain the improved YOLOv8 model; Adopt the weighted box fusion and multi-scale detection fusion strategies to perform weighted average processing on the defect bounding boxes to obtain the target defect bounding box, and at the same time generate the position information of the target defect bounding box; Control the repair device to repair the wood according to the position information of the target defect bounding box and the defect type.

2. The wood repair method according to claim 1, wherein, The processing of the real-time wood image by the large wood defect detection model includes: Use depthwise separable convolution to expand the real-time wood image to obtain two initial branch features; Adopt a sub-module composed of a series of bottleneck structure modules of multiple depthwise separable convolutions to process the initial features in the branch. Each module internally uses two layers of depthwise separable convolutions, including the depthwise separable convolution for feature compression and the depthwise separable convolution for feature expansion. When the number of input channels is the same as the number of output channels, residual connection is adopted. After multiple series of processing, multiple local branch features are formed; Concatenate the two initial branch features and multiple local branch features in the channel dimension; Adopt a lightweight attention module based on convolution to perform adaptive re-weighting on the concatenated features. This module uses two layers of convolution. After the first layer of convolution reduces the dimension and is activated by SiLU, the second layer of convolution increases the dimension and is activated by the Sigmoid activation function to obtain the channel attention weight. Finally, multiply the weight with the original concatenated features channel by channel; Map the features weighted by the adaptive weight through one layer of depthwise separable convolution to the target output channels.

3. The wood repair method according to claim 2, wherein, Before the real-time wood image is expanded by using one layer of depthwise separable convolution, it also includes: Perform global average pooling and max pooling on the real-time wood image by using channel attention, and generate a channel mask through convolution layer mapping; Perform global average pooling and max pooling on the real-time wood image by using spatial attention, and generate a spatial mask through convolution mapping; Fuse the channel mask and the spatial mask through the weighting coefficient α to obtain the output of the initial re-weighting; Input the output of the initial re-weighting into the channel attention and spatial attention again to obtain a new channel mask and a new spatial mask; Fuse the new channel mask and the new spatial mask through the weighting coefficient β to obtain the output of the final re-weighting.

4. The wood repair method according to claim 1, wherein The defect types include cracks and knots.

5. A wood repair device, characterized in that, It includes: A real-time wood image acquisition module for obtaining the real-time image of the wood; The wood defect recognition module is used to input the real-time image of the wood into the large wood defect detection model to obtain the defect bounding box and the defect type. Among them, the large wood defect detection model is based on the improved YOLOv8 model; the Conv module in the backbone network of the initial YOLOv8 model is replaced with the MASFAConv module, the C2f module in the neck network is replaced with the C2f_DSCA module, and the P2 detection head is added to the detection head to obtain the improved YOLOv8 model; The target defect bounding box position information generation module is used to perform weighted average processing on the defect bounding box by adopting the weighted box fusion and multi-scale detection fusion strategies to obtain the target defect bounding box, and simultaneously generate the target defect bounding box position information; The wood repair control module is used to control the repair device to repair the wood according to the target defect bounding box position information and the defect type.

6. A wood repair device, characterized in that Comprising: A memory; And A processor coupled to the memory, the processor being configured to execute the wood repair method according to any one of claims 1 to 4 based on the instructions stored in the memory.

7. A computer-readable storage medium, characterized in that, Stored thereon are computer program instructions which, when executed by the processor, implement the wood repair method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Wood defect detection method and system based on deep neural network

    CN117876663A

  • Wood surface defect detection method, system, medium and equipment

    CN119205758A

  • Wood surface defect detection method based on improved YOLOv81, electronic equipment and storage medium

    CN119478523A

  • Target detection method based on improved yolov8 algorithm

    CN119540542A

  • Wood water paint surface defect detection method based on computer vision

    CN119831958A