Wood surface defect detection method based on uniform complexity prior driving

By constructing a neural network model driven by priors of uniform complexity, the problems of false detection of background texture and loss of structural information of weak defects in wood surface defect detection are solved, and high-precision wood surface defect detection is achieved.

CN122453801APending Publication Date: 2026-07-24SHANDONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG UNIV OF SCI & TECH
Filing Date
2026-05-12
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

In existing technologies for detecting defects on wood surfaces, natural background textures can easily lead to false detections, and structural information of weak defects is easily lost during the multi-scale feature transmission process, resulting in insufficient detection accuracy.

Method used

A neural network model based on unified complexity prior is constructed, and efficient detection of wood surface defects is achieved through the combination of backbone feature extraction, complexity-aware dual-branch routing, deep feature extraction, cross-scale feature fusion and detection head.

Benefits of technology

It improves the accuracy of wood surface defect detection, reduces false detections, maintains the model's lightweight design and meets engineering deployment requirements, and also takes into account the detection of weak defects in complex backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The application discloses a wood surface defect detection method based on unified complexity prior driving and belongs to the technical field of wood surface intelligent detection, which is used for wood surface defect detection and comprises the following steps: taking a wood surface image to be detected as input features, constructing a neural network model based on unified complexity prior driving, performing neural network training, setting a training round threshold, stopping the neural network training when the training round is equal to the training round threshold, and outputting the trained neural network model based on unified complexity prior driving; and inputting the wood surface image to be detected into the trained neural network model based on unified complexity prior driving and outputting a detection result. Through the construction of unified complexity prior, the local texture complexity degree can be introduced into the trunk enhancement and cross-scale fusion process at the same time, so that the complex background suppression and the weak defect are kept in the same condition control framework to realize collaborative optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention discloses a method for detecting defects on wood surfaces based on a priori-driven approach with uniform complexity, belonging to the field of intelligent wood surface detection technology. Background Technology

[0002] In recent years, several enhancement methods have emerged for industrial surface defect detection tasks, such as edge enhancement, frequency domain modeling, attention mechanisms, or cross-scale fusion to improve detection accuracy. These methods have improved the characterization ability of fine-grained defects to some extent, but most of them mainly focus on local optimization in a single stage, such as enhancing local responses only in the backbone network or supplementing high-level semantics only in the feature fusion stage. They lack a unified control mechanism that can simultaneously run through the backbone feature extraction and cross-scale fusion processes.

[0003] The aforementioned problems are particularly prominent for wood surface defect detection tasks. As a natural heterogeneous material, wood's surface typically contains complex textures such as growth rings, fiber orientation, knot edges, and resin flow marks. These natural textures, in terms of local morphology, high-frequency response, and edge structure, are often highly similar to real defects such as cracks, wormholes, knots, and missing knots, easily leading to background-induced false detections. Meanwhile, weak defects such as cracks are usually small in size, elongated in shape, and have low contrast; their identification relies heavily on the local geometric structure in shallow features. Furthermore, during multiple downsampling and semantic aggregation processes, these fine-grained structures are easily covered or weakened by higher-level semantics, resulting in missed detections. Summary of the Invention

[0004] The purpose of this invention is to provide a wood surface defect detection method based on a priori-driven uniform complexity, in order to solve the problems in the prior art where natural background textures are prone to induce false detections and weak defects are prone to structural information loss during multi-scale feature transfer.

[0005] A wood surface defect detection method based on unified complexity priors includes: Using the image of the wood surface to be detected as input features, a neural network model driven by a unified complexity prior is constructed, and the neural network is trained. A training round threshold is set. When the number of training rounds equals the training round threshold, the neural network training is stopped, and the trained neural network model driven by a unified complexity prior is output. The image of the wood surface to be detected is input into a trained neural network model driven by a prior with uniform complexity, and the model outputs the detection results of wood surface defects. The neural network model driven by unified complexity prior includes a backbone feature extraction network and a unified complexity prior construction module. The backbone feature extraction network includes 3 feature extraction units, 2 complexity-aware dual-branch routing modules, a deep feature extraction unit, 2 complexity-guided adaptive fusion modules, a cross-scale feature fusion module, and 3 detection heads.

[0006] The image of the wood surface to be detected is input into the backbone feature extraction network, and then passes through the first feature extraction unit, the second feature extraction unit, the first complexity-aware dual-branch routing module, the third feature extraction unit, the second complexity-aware dual-branch routing module, and the deep feature extraction unit in sequence. The first feature extraction unit outputs the first intermediate layer features. The second feature extraction unit outputs the second intermediate layer features. The third feature extraction unit outputs the third intermediate layer features. ; Will and Input a prior construction module with uniform complexity, perform channel mean projection, and obtain the corresponding grayscale feature map. and Then to and Haar wavelet high-frequency decomposition is performed separately, including three branches: the first branch performs the horizontal high-frequency subband response, the second branch performs the vertical high-frequency subband response, and the third branch performs the diagonal high-frequency subband response. The absolute values ​​of the processing results of the three branches are then added element by element. The space complexity graph is obtained by statistically aggregating the results of element-by-element addition. Statistical aggregation includes upsampling, aggregation, normalization, and truncation. Will Perform global average pooling and output a global complexity scalar. ; Will and Input the first complexity-aware dual-branch routing module and output the enhanced backbone output features. ,Will and Input the second complexity-aware dual-branch routing module and output the enhanced backbone output features. ; Deep feature extraction unit receives Output deep semantic features ,Will , and The second complexity guides the adaptive fusion module, for and Perform channel splicing, and based on Spatial gating is applied to the context semantic injection location, and the output is then fused into a mid-layer. ;Will , and Input the first complexity-guided adaptive fusion module, and... , Perform channel splicing, and based on Modulate the intensity of contextual semantic injection and output a shallow fusion output. ; Will and Input the cross-scale feature fusion module and output the cross-scale fused features.

[0007] Will Input the first detection head; Input the cross-scale fusion features into the second detection head; The cross-scale fusion features and deep semantic features are input into the third detection head; The outputs of the first, second, and third detection heads are used as the detection results for wood surface defects.

[0008] The complexity-aware dual-branch routing module receives input features. , Input channel partitioning units, and then output processing branch features respectively. And identity preserves branch characteristics ; Complexity-aware dual-branch routing module receives Input complexity bias generation unit, output complexity bias : ; In the formula, This is the offset scaling factor; Input the routing feature generation unit, frequency domain enhancement branch unit, and edge enhancement branch unit respectively.

[0009] The routing feature generation unit receives It also extracts routing control features and inputs them into the routing network to obtain the raw routing values. ,Will and Input a dual-branch weight generation unit, output frequency-domain enhanced branch weights and edge enhancement branch weights and to and Set a lower bound and renormalize: ; In the formula, The preset temperature parameters; Frequency domain enhancement branch reception Output frequency domain enhancement features Edge-enhanced branch reception Output edge enhancement features ,Will , , and The common input weighted fusion unit performs weighted fusion to obtain the fused processing branch features. : ; Will and The input channel splicing unit splices the data, and the splicing result is input into the linear projection and residual hybrid unit, which outputs the enhanced backbone output features.

[0010] Frequency domain enhancement branch reception The grayscale projection unit performs mean projection along the channel dimension to obtain the grayscale features for frequency domain analysis. These features are then input into the high-frequency extraction unit to extract the high-frequency response. ,Will Input the global mapping unit, the local mapping unit, and the weighted modulation unit respectively; The processing results from the global mapping unit and the local mapping unit are input together into the gating weight generation unit to generate gating weights. ; Weighted modulation unit receives and ,Will and Element-wise multiplication is performed, and the result is input into the upsampling unit for upsampling. Spatial resolution, upsampling results and The common input residual enhancement unit performs residual superposition, and the output is... .

[0011] Edge Enhanced Branch Reception Input the Sobel-x edge extraction unit, Sobel-y edge extraction unit, Laplacian edge extraction unit, and basic structure branch respectively; The edge responses of the Sobel-x edge extraction unit, the Sobel-y edge extraction unit, and the Laplacian edge extraction unit are as follows: ; ; ; In the formula, The edge response corresponding to the Sobel-x edge extraction unit. The edge response corresponding to the Sobel-y edge extraction unit. This represents the edge response corresponding to the Laplacian edge extraction unit. For Sobel-x operators, For the Sobel-y operator, For Laplacian operators; , and The input is a lightweight gating network, processed by the Softmax function, and the output is the corresponding fusion weights. , , The corresponding weights are input into the weighted fusion unit, and the output is a mixed edge feature. ; Basic structural branch pairs Perform convolution processing to output basic structural features. ,Will and The input feature combination unit is then input into the local spatial attention unit, and the local spatial attention is used to... and The combined features are recalibrated to obtain edge enhancement features. .

[0012] The complexity-guided adaptive fusion module consists of a backbone and prior branches. The backbone architecture of the first and second complexity-guided adaptive fusion modules is the same. The backbone receives the output of at least two neural network layers, which then pass through a channel concatenation unit and a channel compression unit to output a compact representation. , Input the local branch unit and the context branch unit respectively; Local branch unit outputs local features The context branch unit outputs context features. ; The local branch unit is a local branch convolutional structure, and the context branch unit is a context branch large receptive field convolutional structure.

[0013] The first complexity guides the prior branch reception of the adaptive fusion module. Input scalar modulation unit, output scalar modulation coefficient : ; In the formula, This is the lower bound of scalar modulation. For normalization; The main trunk will include contextual features and Element-wise multiplication, the result of element-wise multiplication is Element-by-element addition yields a shallow fusion output. ; The second complexity guides the prior branch reception of the adaptive fusion module. Input spatial gating unit, output spatial gating diagram : ; In the formula, For spatial suppression intensity parameters, This serves as the lower bound of the spatial gating. The main trunk will include contextual features and Element-wise multiplication, the result of element-wise multiplication is Element-by-element addition yields the mid-level fusion output. .

[0014] The global mapping unit includes a global average pooling layer and a 1×1 convolutional layer; the local mapping unit includes a 3×3 convolutional layer and a batch normalization layer; the outputs of the global and local mapping units are fused and then gating weights are generated using a sigmoid activation function; the routing feature generation unit includes a global average pooling layer and two fully connected layers; the high-frequency extraction unit has a fixed convolutional kernel. The wavelet high-frequency filter kernel, Sobel-x operator, Sobel-y operator and Laplacian operator are fixed convolution kernels, and the weight tensor is a preset constant.

[0015] Compared with existing technologies, this invention has the following advantages: By constructing a unified complexity prior, this invention can simultaneously introduce local texture complexity into the backbone enhancement and cross-scale fusion process, enabling complex background suppression and weak defects to achieve synergistic optimization under the same conditional control framework; through complexity-aware dual-branch routing enhancement, a dynamic trade-off can be made between frequency domain enhancement and edge enhancement based on the overall texture complexity of the sample, thereby improving the separability between real defects and natural wood grain; through complexity-guided adaptive fusion, conditional constraints can be applied to contextual semantic injection based on the spatial complexity graph and global complexity scalar, thereby better preserving the local geometric structure upon which weak defects such as cracks depend while introducing necessary high-level semantic information; in the detection stage, the method of this invention does not require the introduction of additional complex post-processing, thus maintaining high detection accuracy while also meeting the requirements of model lightweighting and engineering deployment. Attached Figure Description

[0016] Figure 1This is a schematic diagram of the neural network model structure based on unified complexity prior driving according to the present invention; Figure 2 This is a schematic diagram of the complexity-aware dual-branch routing module of the present invention; Figure 3 This is a schematic diagram of the frequency domain enhancement branch of the present invention; Figure 4 This is a schematic diagram of the edge-enhancing branch structure of the present invention; Figure 5 A schematic diagram of the structure of the first complexity-guided adaptive fusion module; Figure 6 This is a schematic diagram of the structure of the adaptive fusion module, which guides the second level of complexity. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention are described clearly and completely below. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0018] A wood surface defect detection method based on unified complexity priors includes: Using the image of the wood surface to be detected as input features, a neural network model driven by a unified complexity prior is constructed, and the neural network is trained. A training round threshold is set. When the number of training rounds equals the training round threshold, the neural network training is stopped, and the trained neural network model driven by a unified complexity prior is output. The image of the wood surface to be detected is input into a trained neural network model driven by a prior with uniform complexity, and the model outputs the detection results of wood surface defects. The neural network model driven by unified complexity prior includes a backbone feature extraction network and a unified complexity prior construction module. The backbone feature extraction network includes 3 feature extraction units, 2 complexity-aware dual-branch routing modules, a deep feature extraction unit, 2 complexity-guided adaptive fusion modules, a cross-scale feature fusion module, and 3 detection heads.

[0019] The image of the wood surface to be detected is input into the backbone feature extraction network, and then passes through the first feature extraction unit, the second feature extraction unit, the first complexity-aware dual-branch routing module, the third feature extraction unit, the second complexity-aware dual-branch routing module, and the deep feature extraction unit in sequence. The first feature extraction unit outputs the first intermediate layer features. The second feature extraction unit outputs the second intermediate layer features. The third feature extraction unit outputs the third intermediate layer features. ; Will and Input a prior construction module with uniform complexity, perform channel mean projection, and obtain the corresponding grayscale feature map. and Then to and Haar wavelet high-frequency decomposition is performed separately, including three branches: the first branch performs the horizontal high-frequency subband response, the second branch performs the vertical high-frequency subband response, and the third branch performs the diagonal high-frequency subband response. The absolute values ​​of the processing results of the three branches are then added element by element. The space complexity graph is obtained by statistically aggregating the results of element-by-element addition. Statistical aggregation includes upsampling, aggregation, normalization, and truncation. Will Perform global average pooling and output a global complexity scalar. ; Will and Input the first complexity-aware dual-branch routing module and output the enhanced backbone output features. ,Will and Input the second complexity-aware dual-branch routing module and output the enhanced backbone output features. ; Deep feature extraction unit receives Output deep semantic features ,Will , and The second complexity guides the adaptive fusion module, for and Perform channel splicing, and based on Spatial gating is applied to the context semantic injection location, and the output is then fused into a mid-layer. ;Will , and Input the first complexity-guided adaptive fusion module, and... , Perform channel splicing, and based on Modulate the intensity of contextual semantic injection and output a shallow fusion output. ; Will and Input the cross-scale feature fusion module and output the cross-scale fused features.

[0020] Will Input the first detection head; Input the cross-scale fusion features into the second detection head; The cross-scale fusion features and deep semantic features are input into the third detection head; The outputs of the first, second, and third detection heads are used as the detection results for wood surface defects.

[0021] The complexity-aware dual-branch routing module receives input features. , Input channel partitioning units, and then output processing branch features respectively. And identity preserves branch characteristics ; Complexity-aware dual-branch routing module receives Input complexity bias generation unit, output complexity bias : ; In the formula, This is the offset scaling factor; Input the routing feature generation unit, frequency domain enhancement branch unit, and edge enhancement branch unit respectively.

[0022] The routing feature generation unit receives It also extracts routing control features and inputs them into the routing network to obtain the raw routing values. ,Will and Input a dual-branch weight generation unit, output frequency-domain enhanced branch weights and edge enhancement branch weights and to and Set a lower bound and renormalize: ; In the formula, The preset temperature parameters; Frequency domain enhancement branch reception Output frequency domain enhancement features Edge-enhanced branch reception Output edge enhancement features ,Will , , and The common input weighted fusion unit performs weighted fusion to obtain the fused processing branch features. : ; Will and The input channel splicing unit splices the data, and the splicing result is input into the linear projection and residual hybrid unit, which outputs the enhanced backbone output features.

[0023] Frequency domain enhancement branch reception The grayscale projection unit performs mean projection along the channel dimension to obtain the grayscale features for frequency domain analysis. These features are then input into the high-frequency extraction unit to extract the high-frequency response. ,Will Input the global mapping unit, the local mapping unit, and the weighted modulation unit respectively; The processing results from the global mapping unit and the local mapping unit are input together into the gating weight generation unit to generate gating weights. ; Weighted modulation unit receives and ,Will and Element-wise multiplication is performed, and the result is input into the upsampling unit for upsampling. Spatial resolution, upsampling results and The common input residual enhancement unit performs residual superposition, and the output is... .

[0024] Edge Enhanced Branch Reception Input Sobel-x edge extraction units, Sobel-y edge extraction units, Laplacian edge extraction units, and basic structure branches respectively; The edge responses of the Sobel-x edge extraction unit, the Sobel-y edge extraction unit, and the Laplacian edge extraction unit are as follows: ; ; ; In the formula, The edge response corresponding to the Sobel-x edge extraction unit. This represents the edge response corresponding to the Sobel-y edge extraction unit. This represents the edge response corresponding to the Laplacian edge extraction unit. For Sobel-x operators, For the Sobel-y operator, For Laplacian operators; , and The input is a lightweight gating network, processed by the Softmax function, and the output is the corresponding fusion weights. , , The corresponding weights are input into the weighted fusion unit, and the output is a mixed edge feature. ; Basic structural branch pairs Perform convolution processing to output basic structural features. ,Will and The input feature combination unit is then input into the local spatial attention unit, and the local spatial attention is used to... and The combined features are recalibrated to obtain edge enhancement features. .

[0025] The complexity-guided adaptive fusion module consists of a backbone and prior branches. The backbone architecture of the first and second complexity-guided adaptive fusion modules is the same. The backbone receives the output of at least two neural network layers, which then pass through a channel concatenation unit and a channel compression unit to output a compact representation. , Input the local branch unit and the context branch unit respectively; Local branch unit outputs local features The context branch unit outputs context features. ; The local branch unit is a local branch convolutional structure, and the context branch unit is a context branch large receptive field convolutional structure.

[0026] The first complexity guides the prior branch reception of the adaptive fusion module. Input scalar modulation unit, output scalar modulation coefficient : ; In the formula, This is the lower bound of scalar modulation. For normalization; The main trunk will include contextual features and Element-wise multiplication, the result of element-wise multiplication is Element-by-element addition yields a shallow fusion output. ; The second complexity guides the prior branch reception of the adaptive fusion module. Input spatial gating unit, output spatial gating diagram : ; In the formula, For spatial suppression intensity parameters, This serves as the lower bound of the spatial gating. The main trunk will include contextual features and Element-wise multiplication, the result of element-wise multiplication is Element-by-element addition yields the mid-level fusion output. .

[0027] The global mapping unit includes a global average pooling layer and a 1×1 convolutional layer; the local mapping unit includes a 3×3 convolutional layer and a batch normalization layer; the outputs of the global and local mapping units are fused and then gating weights are generated using a sigmoid activation function; the routing feature generation unit includes a global average pooling layer and two fully connected layers; the high-frequency extraction unit has a fixed convolutional kernel. The wavelet high-frequency filter kernel, Sobel-x operator, Sobel-y operator and Laplacian operator are fixed convolution kernels, and the weight tensor is a preset constant.

[0028] The high-frequency extraction unit uses a Haar wavelet high-frequency filter kernel, and the weight tensor includes a horizontal high-frequency filter kernel. Vertical high-frequency filter core and diagonal high-frequency filter core They are respectively: ; ; ; The weight tensor of the Sobel-x operator is: ; The weight tensor of the Sobel-y operator is: ; The weight tensor of the Laplacian operator is: ; The aforementioned fixed convolutional kernels remain unchanged during network training and inference and do not participate in backpropagation updates.

[0029] The following description, in conjunction with the accompanying drawings, further illustrates the present invention's neural network model structure based on a unified complexity prior-driven approach, as shown below. Figure 1 As shown, the input image is sequentially processed through a first feature extraction unit, a second feature extraction unit, a first complexity-aware dual-branch routing module (CADR), a third feature extraction unit, a second complexity-aware dual-branch routing module (CADR), and a deep feature extraction unit; the first feature extraction unit outputs the first intermediate layer features. The second feature extraction unit outputs the second intermediate layer features. The third feature extraction unit outputs the third intermediate layer features. ; and Inputting a uniform complexity prior construct (DOFS) module, after channel mean projection, the responses of three high-frequency subbands are calculated: horizontal, vertical, and diagonal high-frequency subband responses. and The corresponding processing results are added element by element, and the resulting aggregated data generates a space complexity graph. ,right Perform global average pooling to obtain a global complexity scalar. ;Will and The first complexity-aware dual-branch routing module is input, and the processing results are input into the first complexity-guided adaptive fusion module (CGAF) and the third feature extraction unit, respectively. and The processing results are input into the second complexity-aware dual-branch routing module and then into the deep feature extraction unit and the second complexity-guided adaptive fusion module, respectively. The second complexity-guided adaptive fusion module receives the processing results from the second complexity-aware dual-branch routing module, the deep feature extraction unit, and... The processed results are input into the first complexity-guided adaptive fusion module and the cross-scale feature fusion path layer, respectively; the first complexity-guided adaptive fusion module receives the processing results from the first complexity-aware dual-branch routing module, the processing results from the second complexity-guided adaptive fusion module, and... The processed results are input into the first detection head (P3) and the cross-scale feature fusion path layer, respectively. The processing results of the cross-scale feature fusion path layer are divided into two branches. The first branch is directly input into the second detection head (P4), and the second branch and the results of the deep feature extraction unit are input into the third detection head (P5).

[0030] The structure of the complexity-aware dual-branch routing module of this invention is as follows: Figure 2 As shown, input features First, input the channel partitioning unit, then output the processing branch feature. And identity preserves branch characteristics , The routing feature generation unit, frequency domain enhancement branch, and edge enhancement branch are input respectively; the routing feature generation unit receives... It also extracts routing control features and inputs them into the routing network to obtain the raw routing values. Frequency domain enhancement branch reception Output frequency domain enhancement features Edge-enhanced branch reception Output edge enhancement features Complexity-aware dual-branch routing module receives Input complexity bias generation unit, output complexity bias ;Will and Input a dual-branch weight generation unit, output frequency domain enhancement branch weights and edge enhancement branch weights; , The frequency domain enhancement branch weights and edge enhancement branch weights are input together into the weighted fusion unit for weighted fusion to obtain the fused processing branch features. These fused processing branch features are then combined with... The input channel splicing unit splices the data, and the splicing result is input into the linear projection and residual hybrid unit, which outputs the enhanced backbone output features.

[0031] The structure of the frequency domain enhancement branch of this invention is as follows: Figure 3 As shown, frequency domain enhancement branch reception The grayscale projection unit performs mean projection along the channel dimension to obtain the grayscale features for frequency domain analysis. These features are then input into the high-frequency extraction unit to extract the high-frequency response. ,Will The inputs are respectively fed into the global mapping unit, the local mapping unit, and the weighted modulation unit; the processing results of the global mapping unit and the local mapping unit are jointly input into the gating weight generation unit to generate the gating weights. Weighted modulation unit receives and ,Will and Element-wise multiplication is performed, and the result is input into the upsampling unit for upsampling. Spatial resolution, upsampling results and The common input residual enhancement unit performs residual superposition, and the output is... .

[0032] The structure of the edge-enhanced branch of the present invention is as follows: Figure 4 As shown, edge enhancement branch reception The Sobel-x edge extraction unit, Sobel-y edge extraction unit, Laplacian edge extraction unit, and basic structure branch are input respectively. The edge responses corresponding to the Sobel-x, Sobel-y, and Laplacian edge extraction units are then input into a lightweight gating network, processed by the Softmax function, and the corresponding fusion weights are output. , , The corresponding weights are input into the weighted fusion unit, and the output is a mixed edge feature. ; Basic structural branch pairs Perform convolution processing to output basic structural features. ,Will and The input feature combination unit is then input into the local spatial attention unit, and the local spatial attention is used to... and The combined features are recalibrated to obtain edge enhancement features. .

[0033] The structures of the first complexity-guided adaptive fusion module and the second complexity-guided adaptive fusion module of this invention are as follows: Figure 5 and Figure 6 As shown, the first and second complexity-guided adaptive fusion modules have the same backbone architecture. The backbone receives the output of the previous neural network layer, passes it sequentially through the channel concatenation unit and the channel compression unit, and outputs a compact representation. , The local branch unit and the context branch unit are input separately; the local branch unit outputs local features. The context branch unit outputs context features. The local branch unit is a local branch convolutional structure, and the context branch unit is a context branch large receptive field convolutional structure. The first complexity-guided adaptive fusion module receives... Input scalar modulation unit, output scalar modulation coefficient , context features and Element-wise multiplication, the result of element-wise multiplication is Element-by-element addition yields a shallow fusion output. The second complexity guides the adaptive fusion module to receive... Input spatial gating unit, output spatial gating diagram , context features and Element-wise multiplication, the result of element-wise multiplication is Element-by-element addition yields the mid-level fusion output. .

[0034] The following is a further explanation with reference to an embodiment. In this embodiment, the input image is an RGB image of the wood surface to be detected. The image is input into a backbone feature extraction network, and through stepwise downsampling and feature encoding, the following features are obtained: , and , , , , For batch size, For the number of channels, and These represent the spatial dimensions of the corresponding feature maps. Used to characterize shallow local geometric structure information. Used to characterize the texture and structural information of the middle layer. The first and second intermediate layer features are used to represent high-level semantic information. The first and second intermediate layer features are further used to construct a unified complexity prior, and the third intermediate layer features are mainly used for subsequent cross-scale semantic fusion and detection output.

[0035] Construct a unified complexity prior module to generate a space complexity graph based on the features of the first and second intermediate layers. With global complexity scalar This provides a consistent complexity control signal for subsequent backbone enhancement and feature fusion.

[0036] To each and By performing mean projection along the channel dimension, the corresponding grayscale feature map is obtained. and The formula is as follows: ; ; in, This indicates the mean projection operation along the channel dimension.

[0037] Haar wavelet high-frequency decomposition is performed on the grayscale feature map to extract multiple high-frequency sub-band responses. The absolute values ​​of each high-frequency sub-band are statistically aggregated and used as surrogate quantities for the intensity of local texture and structural perturbation. and ; ; ; in, and These represent grayscale feature maps respectively. and In sub-band High frequency response, .

[0038] The two local high-frequency statistical results are upsampled, aggregated, normalized, and truncated to generate a space complexity graph. The formula is as follows: ; in, Indicates an upsampling operation. This indicates a normalization operation. This indicates a truncation operation.

[0039] Space complexity graph Perform global average pooling to obtain a global complexity scalar. The formula is as follows: ; in, This indicates a global average pooling operation.

[0040] Space complexity diagram Used to describe the local complexity at different spatial locations, the global complexity scalar. This describes the overall texture complexity level of the current image or feature. Together, they constitute a unified complexity prior (DOFS), used for unified conditional control of subsequent backbone enhancement and cross-scale fusion.

[0041] A complexity-aware dual-branch routing enhancement module is constructed to dynamically route the frequency domain enhancement branch and the edge enhancement branch based on a unified complexity prior, thereby achieving a balance between complex background suppression and defect structure enhancement. The complexity-aware dual-branch routing enhancement module includes a channel partitioning unit, a frequency domain enhancement branch, an edge enhancement branch, a routing feature generation unit, a routing network, and a fusion output unit. Input features First, the channel partitioning unit is divided along the channel dimension into identity-preserving branch features. and processing branch features ; Processing branch features To obtain frequency domain enhancement features, we will enter the frequency domain enhancement branch and the edge enhancement branch respectively. and edge enhancement features Simultaneously, the routing feature generation unit processes branch features. Extract routing control features and input them into the routing network to obtain the raw routing values. Furthermore, combining this with the global complexity scalar The generated complexity bias The frequency domain enhancement branch weights are obtained. and edge enhancement branch weights Finally, the output unit is used to... and Perform weighted fusion and combine with identity-preserving branch features Combine the features to output the enhanced backbone output characteristics. .

[0042] Input features Divide along the channel dimension into identity-preserving branch features and processing branch features The formula is as follows: ; in, Used to preserve the original representation and stabilize gradient propagation. Used for subsequent enhancement processing.

[0043] Will process branch features By inputting the frequency domain enhancement branch and the edge enhancement branch respectively, the frequency domain enhancement features are obtained. and edge enhancement features .

[0044] In the frequency domain enhancement branch, firstly... Mean projection is performed along the channel dimension to obtain the grayscale features for frequency domain analysis. The high-frequency response was extracted using a fixed high-frequency convolution kernel. The formula is as follows: ; ; in, This represents the fixed high-frequency convolution kernel extraction operator.

[0045] The frequency domain enhancement branch includes grayscale projection units, high-frequency extraction units, global mapping units, local mapping units, and residual enhancement units. Among these, the grayscale projection unit enhances the features of the processing branch. Mean projection is performed along the channel dimension to obtain the grayscale features of single-channel frequency domain analysis. The high-frequency extraction unit uses a fixed high-frequency convolution kernel. Convolution operations are performed to highlight local texture abrupt changes and high-frequency responses to defects. The fixed high-frequency convolution kernel is preferably a 3×3 high-pass convolution kernel, and more preferably a Laplacian high-frequency convolution kernel; in another embodiment, multiple fixed high-frequency convolution kernels can be used in parallel to extract high-frequency responses in different directions, and then the results are summed or concatenated.

[0046] Furthermore, the high-frequency response is input into the global mapping unit and the local mapping unit respectively to generate gating weights. The formula is as follows: ; in, Represents the global mapping unit. Represents a local mapping unit. This represents the Sigmoid activation function.

[0047] The global mapping unit consists of a global average pooling layer, a 1×1 convolutional layer, and a sigmoid activation layer, used to estimate the importance of the current high-frequency response from an overall perspective. The local mapping unit consists of a 3×3 convolutional layer, a batch normalization layer, and a sigmoid activation layer, used to model the spatial distribution of the high-frequency response from a local neighborhood perspective. The outputs of the global and local mapping units are added together to generate the gating weights. .

[0048] Based on gating weight The high-frequency response is weighted and modulated, and then upsampled and superimposed with the input features using residuals to obtain the frequency domain enhanced features. The formula is as follows: ; in, This indicates element-wise multiplication.

[0049] In the edge enhancement branch, the Sobel-x operator, Sobel-y operator, and Laplacian operator are used to process the branch features respectively. Edge extraction is performed to obtain multiple edge responses. The formula is as follows: ; ; ; in, Represents the Sobel-x operator. Represents the Sobel-y operator, This represents the Laplacian operator.

[0050] Multiple edge responses are input into a lightweight gating network to obtain the corresponding fusion weights. , and The formula is as follows: ; in, This indicates a lightweight gating network.

[0051] The lightweight gated network consists of a global average pooling layer, a first 1×1 convolutional layer, a ReLU activation layer, and a second 1×1 convolutional layer, used to process branch features. Generate normalized fusion weights corresponding to multiple edge responses.

[0052] Based on the fusion weights, multiple edge responses are weighted and fused to obtain hybrid edge features. The formula is as follows: ; in, This represents a local spatial attention operation.

[0053] The basic structural branches are preferably convolutional structures consisting of one or more 3×3 convolutional layers, used to preserve the processing branch features. The local spatial attention is preferably composed of a 3×3 convolutional layer and a sigmoid activation layer, used to process mixed edge features. This local spatial attention is used to process the basic texture and structural information related to the target. With basic structural features The combined results are spatially recalibrated to suppress redundant edge responses in the background region and highlight the true defect boundaries.

[0054] In obtaining frequency domain enhancement features and edge enhancement features Then, based on the global complexity scalar Generation complexity bias and compare it with the original routing value output by the routing network. By combining the results, the frequency domain enhancement branch weights are obtained. and edge enhancement branch weights The formula is as follows: ; ; in, This is the offset scaling factor. This refers to the temperature parameter.

[0055] The routing feature generation unit first processes branch features. Global average pooling is performed to obtain a one-dimensional channel description vector; this channel description vector is then input into a routing network consisting of two fully connected layers to output the original routing values ​​with two branches. The first fully connected layer performs dimensionality reduction mapping on the input channel description vector, and the second fully connected layer outputs two components, corresponding to the original response values ​​of the frequency domain enhancement branch and the edge enhancement branch, respectively. A ReLU activation function is preferably set between the two fully connected layers. In another embodiment, the routing network can also adopt a structure of "1×1 convolution + global pooling + fully connected layer," as long as it can be adapted to the processing branch features. Output the original routing values ​​of the two branches .

[0056] Enhanced features in the frequency domain based on branch weights and edge enhancement features Perform weighted fusion to obtain the fused processing branch features. The formula is as follows: ; Among them, when the global complexity scalar When the value is low, increase the frequency domain to enhance the branch weight. When the global complexity scalar When the weight is high, increase the weight of the edge enhancement branch. .

[0057] The fused processing branch features With identity-preserving branch features The channels are concatenated, and the enhanced backbone output features are obtained through linear projection and residual mixing. The formula is as follows: ; in, This indicates a linear projection operation.

[0058] A complexity-guided adaptive fusion module is constructed to perform cross-scale fusion based on the space complexity graph. and global complexity scalar Conditional constraints are applied to context branches to supplement high-level semantics and preserve weak defect-related structures.

[0059] The complexity-guided adaptive fusion module includes a channel compression unit, local branches, context branches, a scalar modulation unit, a spatial gating unit, and a fusion output unit. The channel compression unit is preferably a 1×1 convolutional layer, used to compress the channels of the concatenated features to be fused, resulting in a compact representation. Local branches are preferably composed of two 3×3 convolutional layers to extract and retain local fine-grained structural information; context branches are preferably composed of large-kernel convolutions, dilated convolutions, or depthwise separable convolutions to expand the receptive field and extract high-level contextual semantic information. The scalar modulation unit is scalarized according to the global complexity. Generate scalar modulation coefficients And apply the modulation coefficient to the context branch output. This allows for overall control of the high-level semantic injection strength during the high-resolution shallow fusion stage. The spatial gating unit is based on the spatial complexity graph. Generate position-by-position spatial gating graph And apply the gating graph to the context branch output. This allows for selective suppression of contextual semantic injection at different spatial locations during the mid-level fusion stage. The fusion output unit then fuses the scalar-modulated or spatially gated contextual features with the local branch output element-wise to obtain cross-scale fused features.

[0060] The two features to be fused are concatenated along the channel dimension, and channel compression is performed through convolution to obtain a compact representation. Compact representation Local features are obtained by inputting the local branch and the context branch respectively. and context features The formula is as follows: ; ; in, This represents a locally branched convolutional structure. This represents a context branch large receptive field convolutional structure.

[0061] In the high-resolution shallow fusion stage, based on the global complexity scalar Generate scalar modulation coefficients And utilize the scalar modulation coefficients to define contextual features. Perform global intensity scaling to obtain shallow fusion output. The formula is as follows: ; ; in, This is the lower bound for scalar modulation.

[0062] In the mid-level fusion stage, according to the space complexity diagram Generate position-by-position spatial gating graph And utilize the spatial gating graph to define contextual features. Spatial selective suppression is performed to obtain the mid-layer fused output. The formula is as follows: ; ; in, For spatial suppression intensity parameters, This is the lower bound of the spatial gating.

[0063] The cross-scale feature fusion module of this invention (i.e. Figure 1 The mid-to-span scale feature fusion path includes, when the features to be fused correspond to a high-resolution shallow fusion stage, using shallow fusion output. As the cross-scale fusion feature output; when the feature to be fused corresponds to the intermediate fusion stage, the intermediate fusion output is used. As a cross-scale fusion feature output.

[0064] Through the above design, the complexity-guided adaptive fusion module can dynamically adjust the position and intensity of context semantic injection based on a unified complexity prior, thereby better balancing false detection suppression and weak defect detection in complex wood grain backgrounds.

[0065] After trunk enhancement and cross-scale fusion are completed, the enhanced trunk features and fused cross-scale features are input into the detection head to obtain the wood surface defect detection results, including the classification and location results of the wood surface defects. The wood surface defects can be one or more of the following: cracks, wormholes, live knots, dead knots, missing knots, resin, or crazing. The detection results can be used in wood quality grading, automated sorting, intelligent processing, and online visual inspection systems.

[0066] In another embodiment, the present invention also provides a wood surface defect detection device based on unified complexity prior driving. The device includes an image input module, a feature extraction module, a prior construction module, a complexity-aware dual-branch routing module, a complexity-guided adaptive fusion module, and a result output module.

[0067] The system comprises the following modules: an image input module for acquiring an image of the wood surface to be inspected; a feature extraction module for extracting features from the first, second, and third intermediate layers; a priori construction module for constructing a spatial complexity map and a global complexity scalar; a complexity-aware dual-branch routing module for dynamically weighting the frequency domain enhancement branch and the edge enhancement branch based on the global complexity scalar; a complexity-guided adaptive fusion module for spatial gating and scalar modulation of the context branch based on the spatial complexity map and the global complexity scalar; and a result output module for outputting the category and location information of defects on the wood surface.

[0068] In this embodiment of the invention, three publicly available industrial surface defect datasets are selected, including two wood surface defect detection datasets and one steel surface defect detection dataset. The first surface defect detection dataset originates from a publicly available large-scale wood surface defect image dataset, containing 20,275 images with a resolution of 2800×1024. In this embodiment, approximately 4,000 images are randomly selected to construct samples, and the training, validation, and test sets are divided in a 7:2:1 ratio. Seven typical defect types are retained, including live knots, dead knots, pith, resin, knot cracks, missing knots, and cracks. The second surface defect detection dataset is a publicly available dataset in the field of wood surface defect detection, containing 10 typical wood surface defects. In this embodiment, the training, validation, and test sets are also divided in a 7:2:1 ratio. The steel surface defect detection dataset is a publicly available steel strip surface defect dataset, containing 1800 200×200 images, covering six typical steel surface defects: cracks, inclusions, patches, pitting, rolled-in oxide scale, and scratches. In this embodiment, the training set, validation set, and test set are also divided in a 7:2:1 ratio.

[0069] Precision, Recall, F1-score, and mAP@0.5 were used as evaluation metrics. Training parameters were as follows: the model was implemented using PyTorch and Ultralytics frameworks, and both training and testing were performed on Ubuntu 24.04 using a single NVIDIA RTX 4090 GPU. The basic detection architecture used a modified YOLO11n, initialized with standard pre-trained weights. The optimizer was AdamW, with a total of 150 training epochs, 5 warm-up epochs, a tolerance of 30 early stopping ping epochs, and an initial learning rate of [value missing]. The learning rate decay coefficient is set to The weight decay coefficient is set to The batch size was set to 16. Data augmentation was performed using HSV perturbation, random translation and scaling, horizontal flipping, and Mosaic stitching. The input sizes of the datasets were set to 1280×512, 640×640, and 224×224, respectively.

[0070] The method of this invention was compared with several representative industrial surface defect detection methods, and the results are shown in Tables 1, 2, and 3: Table 1. Comparison of overall performance and detection results for each category of different methods on the first surface defect detection dataset. ; Table 2. Comparison of overall and category detection performance of different methods on the second surface defect detection dataset. ; Table 3. Comparison of overall and category detection performance of different methods on the steel surface defect detection dataset. ; As shown in Table 1, the method of this invention (Ours in the table) achieved the best overall performance on the first surface defect detection dataset, with mAP@0.5, F1-score, Precision, and Recall reaching 83.3%, 79.3%, 81.3%, and 77.4%, respectively. Compared to WDNET-YOLO, the method of this invention improved mAP@0.5, F1-score, and Recall by 4.3, 3.9, and 5.1 percentage points, respectively; compared to I²GF-Net-d, mAP@0.5 was further improved by 4.9 percentage points. These results clearly demonstrate that the method of this invention can more effectively suppress false detections in complex wood grain backgrounds while maintaining weak defect responses. As shown in Table 2, the method of this invention also achieves superior overall performance on the second surface defect detection dataset, with mAP@0.5 reaching 64.6%, and F1-score, Precision, and Recall at 67.0%, 66.2%, and 69.8%, respectively. Compared to I²GF-Net-d, the method of this invention improves mAP@0.5 by 1.0 percentage point and Recall by 2.0 percentage point, indicating that the method of this invention can still maintain good defect detection capability under long-tailed distribution and multi-class imbalance conditions. As shown in Table 3, the method of this invention achieves an mAP@0.5 of 84.6% on the steel surface defect detection dataset, which is the highest result among all compared methods. Compared with I²GF-Net-d, the method of this invention improves mAP@0.5, F1-score, Precision, and Recall by 0.7, 0.5, 0.8, and 0.3 percentage points, respectively. This indicates that the unified complexity prior-driven mechanism proposed in this invention does not depend on the specific texture distribution in the wood scene and has good applicability in the steel surface defect detection task as well.

[0071] Embodiment 2 of this invention further compares model complexity and inference efficiency. Specifically, the complete model of this invention is compared with the baseline model and models with different module configurations in terms of parameter quantity (Params), floating-point operation quantity (GFLOPs), and FPS under single-threaded CPU conditions. The results are shown in Table 4: Table 4. Ablation experiments on the impact of modules on model complexity and detection performance ; As shown in Table 4, after introducing the complexity-aware dual-branch routing module and the complexity-guided adaptive fusion module, the complete model of this invention only experiences limited changes in the number of parameters and computational cost, but significantly improves detection performance and still maintains an inference speed of approximately 17 FPS under single-threaded CPU conditions. This demonstrates that the method of this invention can maintain good model efficiency and engineering deployment potential while ensuring detection accuracy.

[0072] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting wood surface defects based on unified complexity priors, characterized in that, include: Using the image of the wood surface to be detected as input features, a neural network model driven by a unified complexity prior is constructed, and the neural network is trained. A training round threshold is set. When the number of training rounds equals the training round threshold, the neural network training is stopped, and the trained neural network model driven by a unified complexity prior is output. The image of the wood surface to be detected is input into a trained neural network model driven by a prior with uniform complexity, and the model outputs the detection results of wood surface defects. The neural network model driven by unified complexity prior includes a backbone feature extraction network and a unified complexity prior construction module. The backbone feature extraction network includes 3 feature extraction units, 2 complexity-aware dual-branch routing modules, a deep feature extraction unit, 2 complexity-guided adaptive fusion modules, a cross-scale feature fusion module, and 3 detection heads.

2. The wood surface defect detection method based on unified complexity prior driving according to claim 1, characterized in that, The image of the wood surface to be detected is input into the backbone feature extraction network, and then passes through the first feature extraction unit, the second feature extraction unit, the first complexity-aware dual-branch routing module, the third feature extraction unit, the second complexity-aware dual-branch routing module, and the deep feature extraction unit in sequence. The first feature extraction unit outputs the first intermediate layer features. The second feature extraction unit outputs the second intermediate layer features. The third feature extraction unit outputs the third intermediate layer features. ; Will and Input a prior construction module with uniform complexity, perform channel mean projection, and obtain the corresponding grayscale feature map. and Then to and Haar wavelet high-frequency decomposition is performed separately, including three branches: the first branch performs the horizontal high-frequency subband response, the second branch performs the vertical high-frequency subband response, and the third branch performs the diagonal high-frequency subband response. The absolute values ​​of the processing results of the three branches are then added element by element. The space complexity graph is obtained by statistically aggregating the results of element-by-element addition. Statistical aggregation includes upsampling, aggregation, normalization, and truncation. Will Perform global average pooling and output a global complexity scalar. ; Will and Input the first complexity-aware dual-branch routing module and output the enhanced backbone output features. ,Will and Input the second complexity-aware dual-branch routing module and output the enhanced backbone output features. ; Deep feature extraction unit receives Output deep semantic features ,Will , and The second complexity guides the adaptive fusion module, for and Perform channel splicing, and based on Spatial gating is applied to the context semantic injection location, and the output is then fused into a mid-layer. ;Will , and Input the first complexity-guided adaptive fusion module, and... , Perform channel splicing, and based on Modulate the intensity of contextual semantic injection and output a shallow fusion output. ; Will and Input the cross-scale feature fusion module and output the cross-scale fused features.

3. The wood surface defect detection method based on unified complexity prior driving according to claim 2, characterized in that, Will Input the first detection head; Input the cross-scale fusion features into the second detection head; The cross-scale fusion features and deep semantic features are input into the third detection head; The outputs of the first, second, and third detection heads are used as the detection results for wood surface defects.

4. The wood surface defect detection method based on unified complexity prior driving according to claim 3, characterized in that, The complexity-aware dual-branch routing module receives input features. , Input channel partitioning units, and then output processing branch features respectively. And identity preserves branch characteristics ; Complexity-aware dual-branch routing module receives Input complexity bias generation unit, output complexity bias : ; In the formula, This is the offset scaling factor; Input the routing feature generation unit, frequency domain enhancement branch unit, and edge enhancement branch unit respectively.

5. The wood surface defect detection method based on unified complexity prior driving according to claim 4, characterized in that, The routing feature generation unit receives It also extracts routing control features and inputs them into the routing network to obtain the raw routing values. ,Will and Input a dual-branch weight generation unit, output frequency-domain enhanced branch weights and edge enhancement branch weights and to and Set a lower bound and renormalize: ; In the formula, The preset temperature parameters; Frequency domain enhancement branch reception Output frequency domain enhancement features Edge-enhanced branch reception Output edge enhancement features ,Will , , and The common input weighted fusion unit performs weighted fusion to obtain the fused processing branch features. : ; Will and The input channel splicing unit splices the data, and the splicing result is input into the linear projection and residual hybrid unit, which outputs the enhanced backbone output features.

6. The wood surface defect detection method based on unified complexity prior driving according to claim 5, characterized in that, Frequency domain enhancement branch reception The grayscale projection unit performs mean projection along the channel dimension to obtain the grayscale features for frequency domain analysis. These features are then input into the high-frequency extraction unit to extract the high-frequency response. ,Will Input the global mapping unit, the local mapping unit, and the weighted modulation unit respectively; The processing results from the global mapping unit and the local mapping unit are input together into the gating weight generation unit to generate gating weights. ; Weighted modulation unit receives and ,Will and Element-wise multiplication is performed, and the result is input into the upsampling unit for upsampling. Spatial resolution, upsampling results and The common input residual enhancement unit performs residual superposition, and the output is... .

7. The wood surface defect detection method based on unified complexity prior driving according to claim 6, characterized in that, Edge Enhanced Branch Reception Input the Sobel-x edge extraction unit, Sobel-y edge extraction unit, Laplacian edge extraction unit, and basic structure branch respectively; The edge responses of the Sobel-x edge extraction unit, the Sobel-y edge extraction unit, and the Laplacian edge extraction unit are as follows: ; ; ; In the formula, The edge response corresponding to the Sobel-x edge extraction unit. This represents the edge response corresponding to the Sobel-y edge extraction unit. This represents the edge response corresponding to the Laplacian edge extraction unit. For Sobel-x operators, For the Sobel-y operator, For Laplacian operators; , and The input is a lightweight gating network, processed by the Softmax function, and the output is the corresponding fusion weights. , , The corresponding weights are input into the weighted fusion unit, and the output is a mixed edge feature. ; Basic structural branch pairs Perform convolution processing to output basic structural features. ,Will and The input feature combination unit is then input into the local spatial attention unit, and the local spatial attention is used to... and The combined features are recalibrated to obtain edge enhancement features. .

8. The wood surface defect detection method based on unified complexity prior driving according to claim 7, characterized in that, The complexity-guided adaptive fusion module consists of a backbone and prior branches. The backbone architecture of the first and second complexity-guided adaptive fusion modules is the same. The backbone receives the output of at least two neural network layers, which then pass through a channel concatenation unit and a channel compression unit to output a compact representation. , Input the local branch unit and the context branch unit respectively; Local branch unit outputs local features The context branch unit outputs context features. ; The local branch unit is a local branch convolutional structure, and the context branch unit is a context branch large receptive field convolutional structure.

9. The wood surface defect detection method based on unified complexity prior driving according to claim 8, characterized in that, The first complexity guides the prior branch reception of the adaptive fusion module. Input scalar modulation unit, output scalar modulation coefficient : ; In the formula, This is the lower bound of scalar modulation. For normalization; The main trunk will include contextual features and Element-wise multiplication, the result of element-wise multiplication is Element-by-element addition yields a shallow fusion output. ; The second complexity guides the prior branch reception of the adaptive fusion module. Input spatial gating unit, output spatial gating diagram : ; In the formula, For spatial suppression intensity parameters, This serves as the lower bound of the spatial gating. The main trunk will include contextual features and Element-wise multiplication, the result of element-wise multiplication is Element-by-element addition yields the mid-level fusion output. .

10. The wood surface defect detection method based on unified complexity prior driving according to claim 9, characterized in that, The global mapping unit includes a global average pooling layer and a 1×1 convolutional layer; the local mapping unit includes a 3×3 convolutional layer and a batch normalization layer; the outputs of the global mapping unit and the local mapping unit are fused and then gating weights are generated by the Sigmoid activation function. The routing feature generation unit includes a global average pooling layer and two fully connected layers; the high-frequency extraction unit has a fixed convolutional kernel. The wavelet high-frequency filter kernel, Sobel-x operator, Sobel-y operator and Laplacian operator are fixed convolution kernels, and the weight tensor is a preset constant.