A direction-sensitive dual-branch industrial defect detection method and system
Patent Information
- Application Number
- CN202610897937.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-22
- Publication Date
- 2026-09-18
AI Technical Summary
此类卷积在不同空间方向上具有统一的滤波响应,更倾向于学习稳定的类别级语义信息,导致模型在复杂材料纹理背景下,难以有效区分真实的缺陷区域与具有相似统计特性的背景纹理,容易产生分散、断裂甚至扩散的误响应,使得提取的特征判别性不足
[0014] By adopting the above technical solution, the present invention has the following beneficial effects compared with the prior art: The present invention provides a direction-sensitive dual-branch industrial defect detection method and system. After acquiring the surface image of the industrial product to be inspected, the surface image is input into a direction-sensitive dual-branch detection network. This network extracts conventional semantic feature maps and directional texture feature maps through a conventional semantic branch and a directional texture branch, respectively. The directional texture branch explicitly enhances anisotropic texture modeling using a directional bias convolution operator, and retains and aggregates the directional enhancement responses of multiple intermediate stages through a cascaded directional feature aggregation module to generate directional texture feature maps. After fusing the two types of feature maps at key scales, they are input into a multi-path cross-scale fusion neck network for cross-scale feature recombination, generating enhanced multi-scale detection features, which are then input into the detection head to output the defect detection results. The present invention can effectively suppress interference from complex background textures and improve the detection accuracy of directional defects.
Smart Images

Figure CN122780201A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial defect detection technology, and in particular to a direction-sensitive dual-branch industrial defect detection method and system. Background Technology
[0002] In the detection of surface defects in industrial products, defects such as cracks, scratches, and stripe anomalies often appear as elongated linear structures or anisotropic texture anomalies. Their discriminative information highly depends on the orientation, continuity, and structural damage patterns of the local texture. However, existing defect detection methods based on general object detection frameworks often employ isotropic standard convolutions for feature extraction in their backbone networks. These convolutions have a uniform filtering response across different spatial directions and tend to learn stable category-level semantic information. This makes it difficult for the model to effectively distinguish between real defect regions and background textures with similar statistical characteristics in complex material textures, easily producing scattered, broken, or even diffuse false responses, resulting in insufficient discriminative power of the extracted features. Furthermore, even if directional texture cues are initially captured in shallow networks, this fine-grained anomaly information is easily diluted or covered up during continuous downsampling and semantic abstraction in deeper networks, making it difficult to form a continuous and complete representation of the defect structure. Summary of the Invention
[0003] In view of this, the purpose of this invention is to propose a direction-sensitive bi-branch industrial defect detection method and system, which achieves stable detection of anisotropic defects by decoupling the collaborative representation of directional texture and conventional semantics and retaining fine-grained directional response.
[0004] To achieve the aforementioned technical objectives, in a first aspect, the technical solution adopted by the present invention is: a direction-sensitive dual-branch industrial defect detection method, comprising: Acquire a surface image of the industrial product to be inspected. The surface image contains the defect area to be identified and the background texture area. The surface image is input into a preset orientation-sensitive dual-branch detection network, which includes a regular semantic branch and an orientation texture branch. Feature extraction is performed on the surface image using conventional semantic branches to generate a conventional semantic feature map; The surface image is used to extract features by directional texture branch to generate directional texture feature map. The directional texture branch uses directional bias convolution operator to explicitly enhance and model the anisotropic texture pattern and linear defect structure in the surface image. The directional feature aggregation module retains and aggregates the directional enhancement response of multiple intermediate stages to generate a continuously expressed directional texture feature map. By fusing conventional semantic feature maps and directional texture feature maps at preset key scales, a multi-scale joint representation is obtained. The multi-scale joint representation is input into the multi-path cross-scale fusion neck network. The multi-path cross-scale fusion neck network performs cross-scale feature reorganization on the multi-scale joint representation through a top-down semantic injection path and a bottom-up detail backflow path to generate enhanced multi-scale detection features. The enhanced multi-scale detection features are input into the detection head to generate and output the defect detection results of the industrial product to be inspected.
[0005] In some embodiments, feature extraction of the surface image is performed through conventional semantic branches to generate a conventional semantic feature map, including: The surface image is input into the initial convolutional layer of the regular semantic branch. The initial convolutional layer performs preliminary feature mapping on the surface image and generates initial semantic features. The initial semantic features are sequentially passed to multiple stacked semantic bottleneck modules. Each semantic bottleneck module contains several standard convolutional layers and residual connection structures. The standard convolutional layers use isotropic convolutional kernels to abstract and compress the initial semantic features. The residual connection structures add the input and output features of the semantic bottleneck modules element by element to gradually extract structural information with high-level semantic discriminative ability while maintaining stable feature transmission, thereby generating deep semantic features. In the feature transfer process of multiple stacked semantic bottleneck modules, deep semantic features from multiple preset intermediate stages are retained and used as multi-scale semantic candidate features. Multi-scale semantic candidate features are aggregated across levels with the final output features of regular semantic branches to generate a regular semantic feature map containing rich contextual information and spatial details.
[0006] In some embodiments, feature extraction of the surface image is performed through directional texture branching to generate a directional texture feature map, including: The surface image is input into the initial directional convolutional layer of the directional texture branch. The initial directional convolutional layer performs local directional structure perception on the surface image through the directional bias convolution operator to generate initial directional texture features. The initial orientation texture features are input into the cascaded orientation feature aggregation module, which contains a main branch and a reserved branch. The main branch passes the initial orientation texture features to multiple cascaded orientation bottleneck modules in sequence. Each orientation bottleneck module performs orientation enhancement and updates the features through orientation bias convolution operators, generating multiple intermediate-stage orientation enhancement features. The reserved branch performs identity mapping on the initial orientation texture features to generate reserved branch features. The retained branch features, initial directional texture features, and all intermediate stage directional enhancement features are concatenated in the channel dimension to generate multi-stage directional aggregation features. Channel compression convolution is used to reduce the dimensionality and integrate information of multi-stage directional aggregation features, generating a directional texture feature map that continuously expresses the characteristics.
[0007] In some embodiments, initial orientation texture features are input to a cascaded orientation feature aggregation module. This module includes a main branch and a retention branch. The main branch sequentially passes the initial orientation texture features to multiple cascaded orientation bottleneck modules. Each orientation bottleneck module performs orientation enhancement updates on the features using an orientation bias convolution operator, generating multiple intermediate-stage orientation enhancement features, including: Input the initial orientation texture features into the main branch; In the main branch, the initial directional texture features are input to the first directional bottleneck module. The first directional bottleneck module contains a first directional bias convolutional layer and a second directional bias convolutional layer in series. The first directional bias convolutional layer uses an asymmetric convolutional kernel to enhance the main directional structural response of the input features, generating main directional enhanced features. The second directional bias convolutional layer uses an orthogonal directional asymmetric convolutional kernel to perform orthogonal directional context fusion on the main directional enhanced features, generating the first intermediate directional enhanced features. The initial directional texture features and the first intermediate directional enhanced features are then added element-wise through residual connections to output the first directional bottleneck module features. The first direction bottleneck module features are input into the second direction bottleneck module. The second direction bottleneck module contains a third direction bias convolutional layer and a fourth direction bias convolutional layer in series. The third direction bias convolutional layer uses an asymmetric convolutional kernel to enhance the main direction structural response of the first direction bottleneck module features, generating the second main direction enhanced features. The fourth direction bias convolutional layer uses an orthogonal direction asymmetric convolutional kernel to perform orthogonal direction context fusion on the second main direction enhanced features, generating the second intermediate direction enhanced features. The first direction bottleneck module features and the second intermediate direction enhanced features are added element-wise through residual connections to output the second direction bottleneck module features. Repeat the above steps, passing the features output by the bottleneck module in the previous direction to each subsequent bottleneck module in turn. Each bottleneck module is updated with orientation enhancement through the orientation bias convolution operator, and the input features are retained within the module through residual connections, until the feature transfer of all the concatenated bottleneck modules is completed, generating orientation enhancement features for all intermediate stages. The orientation enhancement features for all intermediate stages include the intermediate orientation enhancement features output by each bottleneck module in the direction.
[0008] In some embodiments, conventional semantic feature maps and directional texture feature maps are fused at preset key scales to obtain a multi-scale joint representation, including: Obtain the regular semantic feature map output by the regular semantic branch at a preset key scale, and the directional texture feature map output by the directional texture branch at the same preset key scale; Determine whether the spatial resolution of the conventional semantic feature map and the directional texture feature map are consistent; If the spatial resolution is consistent, the conventional semantic feature map and the directional texture feature map are directly concatenated in the channel dimension to generate the initial dual-branch fusion feature. If the spatial resolutions are inconsistent, an adaptive average pooling operation is used to downsample the higher-resolution feature map in the conventional semantic feature map and the directional texture feature map to the same target size as the lower-resolution feature map. Then, the two are concatenated in the channel dimension to generate the initial dual-branch fusion feature. The initial dual-branch fusion features are reduced in dimensionality and reorganized in terms of information by channel compression convolution, thereby generating dual-branch fusion features at preset key scales. Repeat the above steps to obtain the bi-branch fusion features corresponding to each of the multiple preset key scales, and use all the bi-branch fusion features together as a multi-scale joint representation.
[0009] In some embodiments, the initial bi-branch fusion features are subjected to channel dimensionality reduction and information recombination through preset channel compression convolution to generate bi-branch fusion features at preset key scales, including: The initial bi-branch fusion feature is obtained. The initial bi-branch fusion feature is composed of a regular semantic feature map and a directional texture feature map concatenated in the channel dimension. Its number of channels is the sum of the number of channels of the regular semantic feature map and the number of channels of the directional texture feature map. The initial dual-branch fusion features are input into a preset channel compression convolutional layer. The channel compression convolutional layer uses a convolutional kernel of preset size to perform feature interaction and information extraction in the local spatial neighborhood of the initial dual-branch fusion features to generate intermediate compressed features. By using a channel compression convolutional layer to perform linear transformation and non-linear activation on the channel dimension of the intermediate compressed features, the number of channels of the intermediate compressed features is compressed to the preset target number of channels, generating a compact fused feature after channel compression. The compact fusion features are batch normalized by channel compression convolutional layers, and the numerical distribution of the compact fusion features is normalized and adjusted to generate normalized fusion features. By using channel compression convolutional layers to map the normalized fused features with nonlinear activation functions, nonlinear transformation capabilities are introduced to generate dual-branch fused features with preset key scales.
[0010] In some embodiments, the multi-scale joint representation is input into a multi-path cross-scale fusion neck network. The multi-path cross-scale fusion neck network performs cross-scale feature reorganization on the multi-scale joint representation through a top-down semantic injection path and a bottom-up detail backflow path to generate enhanced multi-scale detection features, including: Obtain multi-scale joint representations, which include high-resolution detail features, mid-level structural features, and high-level semantic features; By using a top-down semantic injection path, high-level semantic features are upsampled to align the spatial resolution of the upsampled high-level semantic features with that of the mid-level structural features, thus obtaining the upsampled high-level semantic features. The upsampled high-level semantic features and mid-level structural features are concatenated through channels, and the concatenation block is used to reorganize the local features of the concatenation result to generate enhanced mid-level features. Upsample the enhanced mid-level features to align the spatial resolution of the upsampled enhanced mid-level features with that of the high-resolution detail features; The upsampled enhanced mid-layer features are concatenated with high-resolution detail features, and the concatenation block is used to reshape the local features of the concatenation result to generate enhanced shallow features. By using a bottom-up detail reflow path, the enhanced shallow features are downsampled to align the spatial resolution of the downsampled enhanced shallow features with that of the enhanced mid-layer features. The downsampled enhanced shallow features and enhanced mid-layer features are concatenated by channels, and the concatenation block is used to reorganize the local features of the concatenation result to generate updated mid-layer features. The updated mid-level features are downsampled to align the spatial resolution of the downsampled updated mid-level features with that of the high-level semantic features. The downsampled updated mid-level features are concatenated with the high-level semantic features, and the concatenation result is locally reorganized through a convolutional block to generate updated high-level features. The enhanced shallow features, updated mid-level features, and updated high-level features are used together as the enhanced multi-scale detection features.
[0011] In some embodiments, high-level semantic features are upsampled through a top-down semantic injection path to align the spatial resolution of the upsampled high-level semantic features with that of the mid-level structural features, resulting in upsampled high-level semantic features, including: We acquire high-level semantic features and mid-level structural features. High-level semantic features have lower spatial resolution and stronger semantic abstraction information, while mid-level structural features have higher spatial resolution and richer structural detail information. Determine the spatial resolution ratio between high-level semantic features and mid-level structural features. The spatial resolution ratio is the ratio obtained by dividing the spatial size of the mid-level structural features by the spatial size of the high-level semantic features. Based on the spatial resolution ratio, the corresponding upsampling mode is selected. The upsampling mode includes one of nearest neighbor interpolation upsampling, bilinear interpolation upsampling, or transposed convolution upsampling. By upsampling, the spatial resolution of high-level semantic features is magnified, so that the spatial size of the magnified high-level semantic features is completely consistent with the spatial size of the mid-level structural features, thus generating spatially aligned high-level semantic features. Spatially aligned high-level semantic features are output as upsampled high-level semantic features.
[0012] In some embodiments, enhanced multi-scale detection features are input to the detection head to generate and output defect detection results for the industrial product to be inspected, including: The enhanced multi-scale detection features are obtained, which include enhanced shallow features, updated middle-level features, and updated high-level features. The enhanced shallow features, updated mid-level features, and updated high-level features are respectively input into the corresponding classification and regression branches in the detection head; The feature response at each preset anchor point is classified and predicted by the classification branch, and the classification confidence of each preset anchor point belongs to the preset defect category is generated. By using a regression branch to predict the bounding box offset of the feature response at each preset anchor point, the predicted bounding box coordinates corresponding to each preset anchor point are generated. Non-maximum suppression is applied to all predicted bounding boxes based on classification confidence to remove redundant predicted bounding boxes with high overlap and low confidence. The predicted bounding boxes retained after screening, their corresponding defect categories, and classification confidence scores are correlated and integrated to generate and output the defect detection results of the industrial products to be inspected.
[0013] In a second aspect, the present invention also provides a direction-sensitive dual-branch industrial defect detection system, applicable to the method described in the first aspect. The system includes an image acquisition module, a direction-sensitive dual-branch detection network module, a key-scale fusion module, a multi-path cross-scale fusion neck network module, and a detection head module. The image acquisition module acquires a surface image of the industrial product to be inspected, the surface image containing a defect region to be identified and a background texture region. The direction-sensitive dual-branch detection network module includes a conventional semantic branch and a directional texture branch. The conventional semantic branch extracts features from the surface image and generates a conventional semantic feature map, while the directional texture branch extracts features from the surface image and generates a directional texture feature map. The directional texture branch internally includes a directional bias convolution operator and a cascaded directional feature aggregation module. The system employs a directional bias convolution operator to explicitly enhance and model anisotropic texture patterns and linear defect structures in surface images. A cascaded directional feature aggregation module preserves and aggregates directional enhancement responses from multiple intermediate stages to generate a continuously expressive directional texture feature map. A key-scale fusion module fuses conventional semantic feature maps and directional texture feature maps at preset key scales to generate a multi-scale joint representation. A multi-path cross-scale fusion neck network module receives the multi-scale joint representation and performs cross-scale feature recombination on the multi-scale joint representation through a top-down semantic injection path and a bottom-up detail backflow path to generate enhanced multi-scale detection features. A detection head module receives the enhanced multi-scale detection features, generates and outputs the defect detection results of the industrial product to be inspected.
[0014] By adopting the above technical solution, the present invention has the following beneficial effects compared with the prior art: The present invention provides a direction-sensitive dual-branch industrial defect detection method and system. After acquiring the surface image of the industrial product to be inspected, the surface image is input into a direction-sensitive dual-branch detection network. This network extracts conventional semantic feature maps and directional texture feature maps through a conventional semantic branch and a directional texture branch, respectively. The directional texture branch explicitly enhances anisotropic texture modeling using a directional bias convolution operator, and retains and aggregates the directional enhancement responses of multiple intermediate stages through a cascaded directional feature aggregation module to generate directional texture feature maps. After fusing the two types of feature maps at key scales, they are input into a multi-path cross-scale fusion neck network for cross-scale feature recombination, generating enhanced multi-scale detection features, which are then input into the detection head to output the defect detection results. The present invention can effectively suppress interference from complex background textures and improve the detection accuracy of directional defects. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a schematic diagram of steps S101 to S107 of the method described in the specific implementation embodiment; Figure 2 This is a schematic diagram of the overall architecture of the direction-sensitive dual-branch detection network described in the specific implementation. Detailed Implementation
[0017] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be particularly noted that the following embodiments are for illustrative purposes only and do not limit the scope of the invention. Similarly, the following embodiments are only some, not all, embodiments of the present invention, and all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Please see Figures 1 to 2 In a first aspect, this embodiment provides a direction-sensitive bi-branch industrial defect detection method, comprising: S101. Obtain a surface image of the industrial product to be inspected. The surface image contains the defect area to be identified and the background texture area. S102. Input the surface image into a preset orientation-sensitive dual-branch detection network, which includes a regular semantic branch and an orientation texture branch. S103. Extract features from the surface image through conventional semantic branches to generate a conventional semantic feature map; S104. The surface image is extracted through the directional texture branch to generate a directional texture feature map. The directional texture branch uses the directional bias convolution operator to explicitly enhance and model the anisotropic texture patterns and linear defect structures in the surface image. The directional texture branch retains and aggregates the directional enhancement responses of multiple intermediate stages through the cascaded directional feature aggregation module to generate a continuously expressed directional texture feature map. S105. The conventional semantic feature map and the directional texture feature map are fused at a preset key scale to obtain a multi-scale joint representation. S106. Input the multi-scale joint representation into the multi-path cross-scale fusion neck network. The multi-path cross-scale fusion neck network performs cross-scale feature reorganization on the multi-scale joint representation through the top-down semantic injection path and the bottom-up detail backflow path to generate enhanced multi-scale detection features. S107. Input the enhanced multi-scale detection features into the detection head to generate and output the defect detection results of the industrial product to be inspected.
[0019] In step S101, the surface image is acquired by an industrial camera deployed on the production line under controlled light. It is stored in the form of a digital matrix, with each pixel encoding the reflectance intensity information of the material surface. In the surface image, the background texture (such as the fiber stripes of bamboo or the rolling marks of steel) itself exhibits a significant directional and periodic structure, while defects (such as cracks and scratches) are local disruptions to the continuity of the background texture. The defect area and the background texture area usually have only slight differences in grayscale, texture direction, or local continuity. Especially in natural materials (such as bamboo and wood), the structural complexity of the background texture itself is the main factor interfering with defect identification.
[0020] In step S102, the training process of the orientation-sensitive dual-branch detection network is based on a dataset containing a large number of surface image samples labeled with defect category and location information. During the training phase, the conventional semantic branch and the orientation texture branch are jointly optimized end-to-end as a whole. The network weights are adjusted through the backpropagation algorithm to minimize the difference between the prediction results and the real annotations. Preferably, before training begins, the conventional semantic branch can be initialized with parameters pre-trained on large-scale general image datasets such as ImageNet to accelerate convergence and improve the extraction capability of semantic features; while the orientation texture branch, due to its special structure, can adopt random initialization or an initialization strategy based on a specific orientation prior. After the surface image is input into the orientation-sensitive dual-branch detection network, the conventional semantic branch and the orientation texture branch can share the same low-level feature extractor at the input layer, or they can be separated at the input stage. The conventional semantic branch focuses on capturing stable macroscopic semantics through isotropic convolution, while the orientation texture branch focuses on perceiving local texture anomalies through orientation-selective operators. The two begin to encode differentiated information from the input side.
[0021] In step S103, the regular semantic branch extracts features by stacking multiple standard convolutional and pooling layers. These standard convolutional layers employ isotropic square kernels, spatially aggregating information without discrimination. Through layer-by-layer abstraction, the receptive field is gradually enlarged, and the low-level edge and color information of the input image is transformed into high-level, category-related semantic features. The regular semantic feature map output by this branch encodes rich semantic category responses in its channel dimension, while preserving coarse location information of the target in its spatial dimension. The network structure parameters of this branch (such as the number of layers, number of channels, and kernel size) are determined during the training phase and remain unchanged during inference.
[0022] In step S104, the directional texture branch filters the input features using a directional bias convolution operator. This directional bias convolution operator has stronger response weights in a specific direction, thus prioritizing the capture of linear structures with continuous orientations in the image, such as cracks or scratches. The cascaded directional feature aggregation module maintains a feature caching and aggregation mechanism. When features are passed between multiple directional enhancement bottleneck modules, it can "copy" or "bypass" the directional structure-sensitive feature responses generated at each intermediate stage. At the end of the module, all retained directional evidence at different levels of abstraction is uniformly combined, preventing fine-grained texture anomaly cues from being diluted during the semanticization process of the deep network. This ensures that the final generated directional texture feature map can continuously and completely express the defect texture information from edges to structures.
[0023] In step S105, the preset key scales are determined by analyzing the true size distribution of defects in the training dataset. Preferably, three levels in the feature pyramid whose receptive field size matches the peak value of the defect size distribution are selected as key scales. The fusion operation, through concatenation along the channel dimension, allows conventional semantic information and directional texture information to coexist in different feature channels at the same spatial location, providing a foundation for subsequent joint inference. This fusion operation is performed independently at each selected key scale, thereby generating a set of multi-scale joint representations.
[0024] In step S106, the top-down semantic injection path first upsamples the highest-level, semantically strongest feature map, for example, by using bilinear interpolation to match its spatial size with the next layer, and then fuses it with the features of the next layer, thereby injecting high-level semantics into the middle layer. The bottom-up detail backflow path downsamples the fused, most detail-rich shallow feature map, for example, by using a convolution with a stride of 2 to match its size with the previous layer, and then fuses it, thereby supplementing the shallow details back into the higher layers. Through this bidirectional, multi-path information interaction, the multi-path cross-scale fusion neck network achieves deep reorganization of multi-scale representations.
[0025] In step S107, the detection head includes parallel classification convolutional layers and regression convolutional layers, which predict each location on the feature map. The classification convolutional layer outputs the probability distribution of the location belonging to a preset defect category, while the regression convolutional layer outputs the bounding box offset of the location relative to the anchor point. These raw prediction values are post-processed, including thresholding and non-maximum suppression, and are finally converted into structured defect detection results containing category, location, and confidence level, and then output.
[0026] This embodiment constructs a dual-branch detection network comprising a conventional semantic branch and a directional texture branch, achieving decoupled learning and collaborative representation of conventional semantic information and direction-sensitive texture information in industrial surface defects. The directional texture branch utilizes directional bias convolution operators and cascaded directional feature aggregation modules to effectively suppress interference from complex background textures and alleviate the attenuation problem of fine-grained directional responses during deep propagation. Simultaneously, through key-scale fusion and multi-path cross-scale fusion neck network, the complementarity and interaction efficiency between features from different levels and sources are enhanced, thereby improving the overall accuracy and robustness of the detection model in recognizing cross-material and multi-morphological defects.
[0027] In some embodiments, feature extraction of the surface image is performed through conventional semantic branches to generate a conventional semantic feature map, including: The surface image is input into the initial convolutional layer of the regular semantic branch. The initial convolutional layer performs preliminary feature mapping on the surface image and generates initial semantic features. The initial semantic features are sequentially passed to multiple stacked semantic bottleneck modules. Each semantic bottleneck module contains several standard convolutional layers and residual connection structures. The standard convolutional layers use isotropic convolutional kernels to abstract and compress the initial semantic features. The residual connection structures add the input and output features of the semantic bottleneck modules element by element to gradually extract structural information with high-level semantic discriminative ability while maintaining stable feature transmission, thereby generating deep semantic features. In the feature transfer process of multiple stacked semantic bottleneck modules, deep semantic features from multiple preset intermediate stages are retained and used as multi-scale semantic candidate features. Multi-scale semantic candidate features are aggregated across levels with the final output features of regular semantic branches to generate a regular semantic feature map containing rich contextual information and spatial details.
[0028] In this embodiment, the initial convolutional layer downsamples and expands the channels of the input image by setting the stride and kernel size, mapping the original pixel space to the feature space. The resulting initial semantic features are the basic data for subsequent hierarchical abstraction.
[0029] The standard convolutional layers within the semantic bottleneck module employ isotropic square convolutional kernels, which perform spatially indiscriminate information mixing and nonlinear transformation of local neighborhoods. By stacking layers one by one, the receptive field is gradually expanded, combining local edges and textures into more discriminative structural responses. The residual connection structure directly superimposes the module's input to the output through identity mapping, thereby alleviating the gradient propagation barrier caused by increased depth and enabling the network to continuously perform high-level abstraction while maintaining the integrity of low-level information.
[0030] Deep semantic features from multiple pre-defined intermediate stages are extracted from stacked bottleneck modules through bypass connections. These features correspond to structural responses at different receptive field scales, and the resulting candidate feature set covers multi-granular information from local details to global contours.
[0031] Cross-level aggregation reorganizes these multi-scale candidate features with the final output features along the channel dimension, so that the final generated conventional semantic feature map retains the spatial details under shallow high resolution and incorporates the category discrimination clues under deep strong semantics.
[0032] This embodiment introduces residual connection structures into the regular semantic branches to ensure the training stability of deep networks. It also utilizes multi-intermediate-stage feature retention and cross-level aggregation mechanisms to ensure that the generated regular semantic feature maps retain shallow spatial details that are crucial for defect localization while maintaining high-level semantic discriminability. This provides a complete semantic foundation for subsequent fusion with directional texture features.
[0033] In some embodiments, feature extraction of the surface image is performed through directional texture branching to generate a directional texture feature map, including: The surface image is input into the initial directional convolutional layer of the directional texture branch. The initial directional convolutional layer performs local directional structure perception on the surface image through the directional bias convolution operator to generate initial directional texture features. The initial orientation texture features are input into the cascaded orientation feature aggregation module, which contains a main branch and a reserved branch. The main branch passes the initial orientation texture features to multiple cascaded orientation bottleneck modules in sequence. Each orientation bottleneck module performs orientation enhancement and updates the features through orientation bias convolution operators, generating multiple intermediate-stage orientation enhancement features. The reserved branch performs identity mapping on the initial orientation texture features to generate reserved branch features. The retained branch features, initial directional texture features, and all intermediate stage directional enhancement features are concatenated in the channel dimension to generate multi-stage directional aggregation features. Channel compression convolution is used to reduce the dimensionality and integrate information of multi-stage directional aggregation features, generating a directional texture feature map that continuously expresses the characteristics.
[0034] In this embodiment, the directional bias convolution operator used in the initial directional convolution layer decomposes the standard square convolution kernel into asymmetric convolution kernels with specific orientations, thereby introducing directional preference into the filtering response, enabling the layer to preferentially perceive local structures with continuous orientations in the image. In some preferred embodiments, the internal calculation process of this directional bias convolution operator can be implemented through the following decomposition modeling: given a surface image input to the initial directional convolution layer... ,in For the real number field, The number of channels in the surface image. The height of the surface image, The width and orientation of the surface image are given by the biased convolution operator on this surface image. The process of directional modeling can be expressed by formula (1): ; In formula (1), This indicates a convolution operation with a kernel size of 3×1, used to enhance the structural response in the principal direction; These are intermediate features enhanced by the structural response in the main direction; This indicates a convolution operation with a kernel size of 1×3, used to fuse contextual information in orthogonal directions; The output is the directional enhancement feature after orthogonal directional context fusion; further, in scenarios where the feature map resolution needs to be reduced, to ensure the consistency of spatial resolution between the directional texture branch and the regular semantic branch, after completing the above main directional convolution operation, supplementary average pooling is applied in the orthogonal direction; finally, the overall forward propagation process of the directional bias convolution operator can be expressed by formula (2): ; In formula (2), This represents the overall mapping consisting of the internal computation process of the directional bias convolution operator described in formula (1) and, if necessary, the supplementary downsampling operation. Represents batch normalization. This represents the SiLU activation function. This is the directional texture feature map that is the final output of the directional bias convolution operator.
[0035] In the cascaded directional feature aggregation module, the main branch updates the features level by level through multiple directional bottleneck modules connected in series. Within each directional bottleneck module, a directional bias convolution operator extracts the structural response along a specific direction, and residual connections maintain the stability of information transmission, thereby generating directional enhancement features at different levels of abstraction. The retention branch performs an identity mapping on the input initial directional texture features, thus retaining a copy of the original directional response at the end of the module, untouched by subsequent convolutional updates.
[0036] The cascading operation along the channel dimension preserves the branch features, initial directional texture features, and all intermediate stage directional enhancement features, and splices them along the channel direction, so that directional responses of different abstraction levels and different update stages coexist in the same feature tensor, thus forming an information-rich multi-stage directional aggregation feature.
[0037] Channel compression convolution performs local neighborhood feature interaction and channel dimensionality reduction on multi-stage directional aggregation features through convolution kernels of preset size, compressing the number of channels after cascading expansion to the preset target number of channels, eliminating redundancy while preserving multi-stage directional information, and generating a compact and continuously expressive directional texture feature map.
[0038] This embodiment, through the parallel design of the main branch and the reserved branch, enables the directional texture branch to retain the original directional evidence without modification and accumulate the deep directional response enhanced through multiple stages during the process of progressive abstraction. Then, through channel cascading and compression, it achieves efficient integration of multi-level directional information, thereby effectively alleviating the problem of attenuation of directional sensitive features in deep propagation and providing a complete directional texture representation for subsequent fusion with conventional semantic features.
[0039] In some embodiments, initial orientation texture features are input to a cascaded orientation feature aggregation module. This module includes a main branch and a retention branch. The main branch sequentially passes the initial orientation texture features to multiple cascaded orientation bottleneck modules. Each orientation bottleneck module performs orientation enhancement updates on the features using an orientation bias convolution operator, generating multiple intermediate-stage orientation enhancement features, including: Input the initial orientation texture features into the main branch; In the main branch, the initial directional texture features are input to the first directional bottleneck module. The first directional bottleneck module contains a first directional bias convolutional layer and a second directional bias convolutional layer in series. The first directional bias convolutional layer uses an asymmetric convolutional kernel to enhance the main directional structural response of the input features, generating main directional enhanced features. The second directional bias convolutional layer uses an orthogonal directional asymmetric convolutional kernel to perform orthogonal directional context fusion on the main directional enhanced features, generating the first intermediate directional enhanced features. The initial directional texture features and the first intermediate directional enhanced features are then added element-wise through residual connections to output the first directional bottleneck module features. The first direction bottleneck module features are input into the second direction bottleneck module. The second direction bottleneck module contains a third direction bias convolutional layer and a fourth direction bias convolutional layer in series. The third direction bias convolutional layer uses an asymmetric convolutional kernel to enhance the main direction structural response of the first direction bottleneck module features, generating the second main direction enhanced features. The fourth direction bias convolutional layer uses an orthogonal direction asymmetric convolutional kernel to perform orthogonal direction context fusion on the second main direction enhanced features, generating the second intermediate direction enhanced features. The first direction bottleneck module features and the second intermediate direction enhanced features are added element-wise through residual connections to output the second direction bottleneck module features. Repeat the above steps, passing the features output by the bottleneck module in the previous direction to each subsequent bottleneck module in turn. Each bottleneck module is updated with orientation enhancement through the orientation bias convolution operator, and the input features are retained within the module through residual connections, until the feature transfer of all the concatenated bottleneck modules is completed, generating orientation enhancement features for all intermediate stages. The orientation enhancement features for all intermediate stages include the intermediate orientation enhancement features output by each bottleneck module in the direction.
[0040] In this embodiment, the asymmetric convolutional kernel used in the first directional biased convolutional layer has a larger kernel coverage area along the main direction than in the orthogonal direction. This non-square receptive field design enables the layer to have a preference for extracting structural responses along a specific direction during the parameter initialization stage. The orthogonal asymmetric convolutional kernel used in the second directional biased convolutional layer extends its kernel coverage area along a direction orthogonal to the first layer, thereby supplementing the activated structural responses in the main direction enhancement features with neighborhood context in the vertical direction. The residual connection adds the initial directional texture features to the first intermediate direction enhancement features element-wise. This operation does not introduce additional learnable parameters and ensures that the original structural information at the module input is still available at the output through identity mapping, thereby avoiding gradient propagation barriers caused by depth stacking.
[0041] The third and fourth biased convolutional layers perform the same functional roles as the first and second biased convolutional layers, respectively. The difference lies in that the input features already carry the accumulated directional enhancement responses from the previous stage; therefore, the subsequent asymmetric convolutional operations further strengthen the already enhanced structure at a higher level. The residual connection adds the first-direction bottleneck module features and the second intermediate-direction enhancement features element-wise, thereby maintaining the integrity of the directional responses output from the previous stage in subsequent processing.
[0042] The above processing steps are repeated sequentially in a cascaded order. Each directional bottleneck module uses the output of the previous module as its input. The convolution kernel parameters used by the directional bias convolution operators within each module are automatically adjusted based on the data distribution during the training phase using the backpropagation algorithm. The residual connections within each module always maintain a direct transmission path for the input features. After all the cascaded directional bottleneck modules have completed feature transmission, the generated directional enhancement features for all intermediate stages include the intermediate directional enhancement features output by each directional bottleneck module. These features correspond to the directional structural responses at different processing stages. Their receptive fields expand progressively with module stacking, and the encoded directional information gradually transitions from local edges to global structural continuity.
[0043] This embodiment uses decomposition modeling of asymmetric convolution kernels to enable the directional bias convolution operator to extract structural responses step by step along the principal and orthogonal directions. The availability of input features at each stage at the module output is maintained through residual connections. Then, the directional enhancement effect is accumulated at different levels of abstraction through serial reuse, thereby establishing a stable perception capability for continuous linear structures and anisotropic textures in deep features.
[0044] In some embodiments, conventional semantic feature maps and directional texture feature maps are fused at preset key scales to obtain a multi-scale joint representation, including: Obtain the regular semantic feature map output by the regular semantic branch at a preset key scale, and the directional texture feature map output by the directional texture branch at the same preset key scale; Determine whether the spatial resolution of the conventional semantic feature map and the directional texture feature map are consistent; If the spatial resolution is consistent, the conventional semantic feature map and the directional texture feature map are directly concatenated in the channel dimension to generate the initial dual-branch fusion feature. If the spatial resolutions are inconsistent, an adaptive average pooling operation is used to downsample the higher-resolution feature map in the conventional semantic feature map and the directional texture feature map to the same target size as the lower-resolution feature map. Then, the two are concatenated in the channel dimension to generate the initial dual-branch fusion feature. The initial dual-branch fusion features are reduced in dimensionality and reorganized in terms of information by channel compression convolution, thereby generating dual-branch fusion features at preset key scales. Repeat the above steps to obtain the bi-branch fusion features corresponding to each of the multiple preset key scales, and use all the bi-branch fusion features together as a multi-scale joint representation.
[0045] In this embodiment, the preset key scale is determined by analyzing the size distribution of defect instances in the training dataset. The selected scale usually corresponds to the level in the feature pyramid where the receptive field size matches the peak value of the defect size distribution, thereby ensuring that the extracted conventional semantic feature map and directional texture feature map can cover the typical scale range of the target defect in terms of spatial resolution.
[0046] The criterion for determining spatial resolution consistency is whether the height and width values of the two feature maps are exactly equal. If the values are equal, the conventional semantic feature map and the directional texture feature map are already aligned pixel-by-pixel in the spatial dimension and can be directly stitched together in the channel dimension. If the values are not equal, an adaptive average pooling operation is used to downsample the higher-resolution feature map to the same target size as the lower-resolution feature map. The output size of this pooling operation is directly specified by the target size, and the pooling window size is automatically calculated based on the ratio of the input size to the output size, thus achieving spatial alignment without manually setting window parameters.
[0047] The concatenation operation along the channel axis merges two feature maps, resulting in a feature map with the number of channels equal to the sum of the number of channels in the two original feature maps, while maintaining the same spatial resolution. The generated initial bi-branch fused feature simultaneously encodes both conventional semantic responses and directional texture responses in different channels at the same spatial location.
[0048] Channel compression convolution uses a convolutional kernel of a preset size to perform feature interaction within the local spatial neighborhood of the initial bi-branch fused features, while compressing the number of channels after concatenation and dilation to a preset target number of channels. The number of output channels of this convolutional layer is fixed before training. The compressed features eliminate channel redundancy while retaining conventional semantic information and directional texture information. The generated bi-branch fused features are the fusion results at the preset key scale.
[0049] The above processing steps are applied one by one to each preset key scale. Spatial alignment, channel stitching, and channel compression operations are performed independently at each scale. The final output of all bi-branch fused features constitutes a multi-scale joint representation. This multi-scale joint representation encodes the conventional semantic information and directional texture information of defects at different scales at different spatial resolutions.
[0050] This embodiment ensures the spatial consistency of dual-branch features before fusion by using spatial resolution judgment and adaptive pooling alignment. It achieves co-location storage of the two types of information through channel splicing and eliminates redundancy through channel compression. This establishes a collaborative representation of conventional semantics and directional texture at multiple key scales, providing a complete and scale-aligned input foundation for subsequent cross-scale fusion of the neck network.
[0051] In some embodiments, the initial bi-branch fusion features are subjected to channel dimensionality reduction and information recombination through preset channel compression convolution to generate bi-branch fusion features at preset key scales, including: The initial bi-branch fusion feature is obtained. The initial bi-branch fusion feature is composed of a regular semantic feature map and a directional texture feature map concatenated in the channel dimension. Its number of channels is the sum of the number of channels of the regular semantic feature map and the number of channels of the directional texture feature map. The initial dual-branch fusion features are input into a preset channel compression convolutional layer. The channel compression convolutional layer uses a convolutional kernel of preset size to perform feature interaction and information extraction in the local spatial neighborhood of the initial dual-branch fusion features to generate intermediate compressed features. By using a channel compression convolutional layer to perform linear transformation and non-linear activation on the channel dimension of the intermediate compressed features, the number of channels of the intermediate compressed features is compressed to the preset target number of channels, generating a compact fused feature after channel compression. The compact fusion features are batch normalized by channel compression convolutional layers, and the numerical distribution of the compact fusion features is normalized and adjusted to generate normalized fusion features. By using channel compression convolutional layers to map the normalized fused features with nonlinear activation functions, nonlinear transformation capabilities are introduced to generate dual-branch fused features with preset key scales.
[0052] In this embodiment, the number of channels of the initial dual-branch fusion feature is obtained by directly adding the number of channels of the conventional semantic feature map and the number of channels of the directional texture feature map. This value is determined by the number of channels output by each of the two branches before training, and the splicing operation itself does not introduce additional parameters.
[0053] The pre-defined convolutional kernel size used in the pre-defined channel compression convolutional layer determines the breadth of feature interaction within the local spatial neighborhood through its spatial coverage. The kernel slides over the initial bi-branch fused features with a pre-defined stride, covering a local spatial region each time it slides. It then performs a weighted summation of the feature values across all channels within that region, thereby achieving information mixing within the spatial neighborhood. The generated intermediate compressed features retain the layout structure of the initial bi-branch fused features in the spatial dimension, while the number of channels remains unchanged in the channel dimension.
[0054] A linear transformation along the channel dimension maps intermediate compressed features linearly using a learnable weight matrix. The number of rows in this weight matrix corresponds to the preset target number of channels, and the number of columns corresponds to the number of channels in the intermediate compressed features. Matrix multiplication compresses the input channel number to the target channel number. A non-linear activation function introduces non-linear transformation capability after each linear transformation, enabling the channel compression convolutional layer to fit non-linear mapping relationships. The resulting compact fused features, after channel compression, are reduced to the preset target number of channels while maintaining spatial resolution.
[0055] Batch normalization normalizes the compact fused features along the channel dimension. Specifically, it calculates the mean and variance of each channel along both the batch and spatial dimensions. These mean and variance are then used to center and scale the compact fused features, ensuring stable mean and variance across all channels. This accelerates training convergence and mitigates internal covariate shift issues. The normalized fused features have a standardized numerical distribution.
[0056] The nonlinear activation function mapping applies a nonlinear transformation to the normalized fusion features element by element. By introducing the nonlinear transformation capability, the channel compression convolutional layer can express more complex feature mapping relationships. The generated bi-branch fusion features at the preset key scale are the final output fusion results at that scale.
[0057] This embodiment achieves local spatial interaction through a pre-sized convolutional kernel, uses learnable linear transformation and nonlinear activation to achieve channel dimensionality reduction and nonlinear mapping, and then stabilizes the numerical distribution through batch normalization. This eliminates channel redundancy while ensuring the integrity of the fused feature information, making the generated bi-branch fused features compact and well-distributed in the channel dimension, providing a scale-aligned and dimensionally unified feature foundation for the input of the subsequent cross-scale fusion neck network.
[0058] In some embodiments, the multi-scale joint representation is input into a multi-path cross-scale fusion neck network. The multi-path cross-scale fusion neck network performs cross-scale feature reorganization on the multi-scale joint representation through a top-down semantic injection path and a bottom-up detail backflow path to generate enhanced multi-scale detection features, including: Obtain multi-scale joint representations, which include high-resolution detail features, mid-level structural features, and high-level semantic features; By using a top-down semantic injection path, high-level semantic features are upsampled to align the spatial resolution of the upsampled high-level semantic features with that of the mid-level structural features, thus obtaining the upsampled high-level semantic features. The upsampled high-level semantic features and mid-level structural features are concatenated through channels, and the concatenation block is used to reorganize the local features of the concatenation result to generate enhanced mid-level features. Upsample the enhanced mid-level features to align the spatial resolution of the upsampled enhanced mid-level features with that of the high-resolution detail features; The upsampled enhanced mid-layer features are concatenated with high-resolution detail features, and the concatenation block is used to reshape the local features of the concatenation result to generate enhanced shallow features. By using a bottom-up detail reflow path, the enhanced shallow features are downsampled to align the spatial resolution of the downsampled enhanced shallow features with that of the enhanced mid-layer features. The downsampled enhanced shallow features and enhanced mid-layer features are concatenated by channels, and the concatenation block is used to reorganize the local features of the concatenation result to generate updated mid-layer features. The updated mid-level features are downsampled to align the spatial resolution of the downsampled updated mid-level features with that of the high-level semantic features. The downsampled updated mid-level features are concatenated with the high-level semantic features, and the concatenation result is locally reorganized through a convolutional block to generate updated high-level features. The enhanced shallow features, updated mid-level features, and updated high-level features are used together as the enhanced multi-scale detection features.
[0059] In this embodiment, high-resolution detail features, mid-level structural features, and high-level semantic features correspond to outputs with different downsampling ratios in the feature pyramid (the specific ratio values are determined based on the total downsampling ratio of the orientation-sensitive dual-branch detection network and the position of the intermediate output node). The spatial resolution of the three decreases sequentially, while the number of channels increases sequentially, thereby encoding multi-granular information from local texture to global semantics.
[0060] The top-down semantic injection path starts with high-level semantic features, upsampling them to match the spatial size of mid-level structural features. The amplified high-level semantic features and mid-level structural features are spatially aligned pixel-by-pixel. Channel concatenation merges them along the channel axis, resulting in a feature map with the sum of the channels of both features, achieving the same spatial resolution as the mid-level structural features. Convolutional blocks perform local feature reshaping on the concatenated result. Through stacked convolutional layers and non-linear activation functions, the multi-source information after concatenation is mixed and filtered within the local spatial neighborhood. The resulting enhanced mid-level features retain mid-level structural details while incorporating high-level semantic guidance.
[0061] Subsequently, the enhanced mid-layer features are upsampled again to enlarge their spatial size to match the high-resolution detail features. Channel splicing and convolution block reshaping operations are performed sequentially. The resulting enhanced shallow layer features retain high-resolution spatial details while gaining semantic support from mid-layer and high-layer features.
[0062] The bottom-up detail reflow path starts by enhancing shallow features. Downsampling is used to reduce their spatial size to match that of the enhanced mid-level features. After channel concatenation and convolutional block reshaping, the generated updated mid-level features supplement the mid-level features with the texture details preserved in the shallow layers. Further downsampling is applied to the updated mid-level features to reduce their spatial size to match that of the high-level semantic features. After channel concatenation and convolutional block reshaping, the generated updated high-level features inject mid- and low-level detail information into the high-level semantic features.
[0063] The enhanced shallow features, updated mid-level features, and updated high-level features together constitute the enhanced multi-scale detection features. The three correspond to high, medium, and low spatial resolution levels, respectively. The features of each level have obtained cross-scale information supplementation through semantic injection or detail backflow path.
[0064] This embodiment uses a top-down approach to inject high-level semantics into mid-to-low-level features to enhance the semantic guidance capability of shallow features, and uses a bottom-up approach to transmit shallow details back to mid-to-high-level features to compensate for the spatial detail loss of high-level features. This enables bidirectional information collaboration between multi-scale features before detection, so that the generated enhanced multi-scale detection features have a more complete texture and semantic joint expression capability at each level.
[0065] In some embodiments, high-level semantic features are upsampled through a top-down semantic injection path to align the spatial resolution of the upsampled high-level semantic features with that of the mid-level structural features, resulting in upsampled high-level semantic features, including: We acquire high-level semantic features and mid-level structural features. High-level semantic features have lower spatial resolution and stronger semantic abstraction information, while mid-level structural features have higher spatial resolution and richer structural detail information. Determine the spatial resolution ratio between high-level semantic features and mid-level structural features. The spatial resolution ratio is the ratio obtained by dividing the spatial size of the mid-level structural features by the spatial size of the high-level semantic features. Based on the spatial resolution ratio, the corresponding upsampling mode is selected. The upsampling mode includes one of nearest neighbor interpolation upsampling, bilinear interpolation upsampling, or transposed convolution upsampling. By upsampling, the spatial resolution of high-level semantic features is magnified, so that the spatial size of the magnified high-level semantic features is completely consistent with the spatial size of the mid-level structural features, thus generating spatially aligned high-level semantic features. Spatially aligned high-level semantic features are output as upsampled high-level semantic features.
[0066] In this embodiment, the spatial resolution ratio between high-level semantic features and mid-level structural features is obtained by directly calculating the quotient of their spatial dimensions. This quotient is the magnification factor required for the upsampling operation. When the magnification factor is an integer, the upsampling operation can achieve pixel filling through regular interpolation; when the magnification factor is a non-integer, grayscale estimation of the target position is required through an interpolation algorithm.
[0067] When selecting the appropriate upsampling mode based on the spatial resolution ratio, nearest-neighbor interpolation upsampling fills the new pixel position by copying the value of the nearest pixel, resulting in the lowest computational cost, but the upsampled high-level semantic features may have block artifacts; bilinear interpolation upsampling calculates the new pixel value by weighted averaging of the four pixels surrounding the target position, resulting in smooth transitions in the upsampled high-level semantic features with moderate computational cost; transposed convolutional upsampling upsamples high-level semantic features using learnable convolution kernel parameters, but the quality of the spatially aligned high-level semantic features generated is affected by the training data and introduces additional learnable parameters. The choice of the three upsampling modes is based on the trade-off between detection accuracy and computational cost.
[0068] When spatially upsampling high-level semantic features using the selected upsampling mode, the upsampling operation only changes the spatial size of the feature map, not the number of channels. The upscaled high-level semantic features and the mid-level structural features are exactly equal in height and width, which means they are spatially aligned high-level semantic features.
[0069] This embodiment determines the upsampling ratio by calculating the spatial resolution ratio, selects the corresponding upsampling mode based on the numerical characteristics of the ratio and the accuracy requirements, and then amplifies the high-level semantic features to a spatial size consistent with the mid-level structural features through the selected mode. This enables the efficient transfer of high-level semantic information to mid-level features in the top-down semantic injection path, providing a spatially aligned input basis for subsequent channel splicing and local feature reorganization.
[0070] In some embodiments, enhanced multi-scale detection features are input to the detection head to generate and output defect detection results for the industrial product to be inspected, including: The enhanced multi-scale detection features are obtained, which include enhanced shallow features, updated middle-level features, and updated high-level features. The enhanced shallow features, updated mid-level features, and updated high-level features are respectively input into the corresponding classification and regression branches in the detection head; The feature response at each preset anchor point is classified and predicted by the classification branch, and the classification confidence of each preset anchor point belongs to the preset defect category is generated. By using a regression branch to predict the bounding box offset of the feature response at each preset anchor point, the predicted bounding box coordinates corresponding to each preset anchor point are generated. Non-maximum suppression is applied to all predicted bounding boxes based on classification confidence to remove redundant predicted bounding boxes with high overlap and low confidence. The predicted bounding boxes retained after screening, their corresponding defect categories, and classification confidence scores are correlated and integrated to generate and output the defect detection results of the industrial products to be inspected.
[0071] In this embodiment, enhancing shallow features, updating mid-level features, and updating high-level features correspond to high, mid, and low spatial resolution levels, respectively. The feature map of each level is independently fed into the corresponding classification and regression branches in the detection head. The classification and regression branches share the same enhanced shallow features, updated mid-level features, or updated high-level features (specifically, the features of the current processing level), but each maintains independent learnable parameters, thereby decoupling the classification task from the localization task.
[0072] When the classification branch performs classification prediction on the feature response at each preset anchor point, the feature map is divided into equally spaced sampling points using a predefined grid, with each sampling point corresponding to a spatial location on the feature map. The classification branch uses convolution operations to map the feature response at each preset anchor point into a probability distribution with a dimension equal to the preset number of defect categories. The generated classification confidence score is the probability value that the location belongs to each type of defect.
[0073] When the regression branch predicts the bounding box offset for the feature responses at each preset anchor point, the predicted bounding box offset typically contains four components, corresponding to the horizontal coordinate offset of the predicted box relative to the center point of the preset anchor point, the vertical coordinate offset of the center point, the width scaling ratio, and the height scaling ratio. The regression branch maps the feature responses at each preset anchor point to four offset values through a convolution operation, and the generated predicted bounding box coordinates are calculated by applying the offsets to the initial coordinates of the preset anchor point.
[0074] Non-maximum suppression (NMS) filtering uses classification confidence as the ranking criterion. First, the predicted bounding box with the highest classification confidence is selected as the retained box. Then, the intersection-union ratio (IUR) of this retained box with all other predicted bounding boxes is calculated, and redundant predicted bounding boxes with IUR exceeding a preset threshold are removed. This process is repeated until all predicted bounding boxes have been processed. The preset threshold is usually determined before training based on performance on the validation set, and is used to balance recall and precision.
[0075] The predicted bounding boxes retained after filtering, along with their corresponding defect categories and classification confidence scores, are associated and integrated. The resulting defect detection results are typically output in a standardized format, including the category label, bounding box coordinates, and confidence score for each detected defect instance.
[0076] This embodiment achieves independent prediction of defect category and location through the decoupling design of classification and regression branches. It establishes the correspondence between feature response and spatial location through a preset anchor point mechanism, and eliminates redundant detection through non-maximum suppression screening. This transforms the enhanced multi-scale detection features into structured, readable and low-redundancy defect detection results, providing a direct basis for subsequent defect location and classification decisions in industrial settings.
[0077] In a second aspect, this embodiment also provides a direction-sensitive dual-branch industrial defect detection system, applicable to the method described in the first aspect. The system includes an image acquisition module, a direction-sensitive dual-branch detection network module, a key-scale fusion module, a multi-path cross-scale fusion neck network module, and a detection head module. The image acquisition module is used to acquire a surface image of the industrial product to be inspected, the surface image containing a defect region to be identified and a background texture region. The direction-sensitive dual-branch detection network module includes a conventional semantic branch and a directional texture branch. The conventional semantic branch is used to extract features from the surface image and generate a conventional semantic feature map, and the directional texture branch is used to extract features from the surface image and generate a directional texture feature map. The directional texture branch is internally configured with a directional bias convolution operator and a cascaded directional feature aggregation module. The system employs a directional bias convolution operator to explicitly enhance and model anisotropic texture patterns and linear defect structures in surface images. A cascaded directional feature aggregation module preserves and aggregates directional enhancement responses from multiple intermediate stages to generate a continuously expressive directional texture feature map. A key-scale fusion module fuses conventional semantic feature maps and directional texture feature maps at preset key scales to generate a multi-scale joint representation. A multi-path cross-scale fusion neck network module receives the multi-scale joint representation and performs cross-scale feature recombination on the multi-scale joint representation through a top-down semantic injection path and a bottom-up detail backflow path to generate enhanced multi-scale detection features. A detection head module receives the enhanced multi-scale detection features, generates and outputs the defect detection results of the industrial product to be inspected.
[0078] In this embodiment, the image acquisition module is deployed at the production line workstation, and its output is connected to the input of the orientation-sensitive dual-branch detection network module. The orientation-sensitive dual-branch detection network module internally processes the acquired surface image in parallel using a conventional semantic branch and an orientation texture branch. The key-scale fusion module receives the outputs of the two branches and performs fusion at a preset key scale. The multi-path cross-scale fusion neck network module performs cross-scale recombination of the multi-scale joint representation through bidirectional paths from top to bottom and bottom to top. The detection head module receives the enhanced multi-scale detection features, generates defect detection results, and outputs them. This system, through pipelined data transfer and functional collaboration between modules, maps the orientation-sensitive dual-branch detection network's method architecture into a hardware-deployable modular system, achieving fully automated processing from image acquisition to detection result output.
[0079] By adopting the above technical solutions, this invention differs from existing technologies and has the following beneficial effects: By constructing a direction-sensitive dual-branch detection network containing a conventional semantic branch and a directional texture branch, it achieves decoupled learning and collaborative representation of conventional semantic information and direction-sensitive texture information in industrial surface images. The directional texture branch uses a directional bias convolution operator to explicitly enhance and model anisotropic texture patterns and linear defect structures, and retains and aggregates the directional enhancement responses of multiple intermediate stages through a cascaded directional feature aggregation module, effectively alleviating the attenuation problem of directional sensitive features in deep propagation; after the conventional semantic feature map and the directional texture feature map are fused at a preset key scale to form a multi-scale joint representation, it is then reorganized across scales through a multi-path cross-scale fusion neck network via a top-down semantic injection path and a bottom-up detail backflow path, so that features at each level simultaneously receive high-level semantic guidance and shallow detail supplementation, enhancing the complementarity and interaction efficiency between features from different levels and sources; finally, the detection head outputs detection results containing defect category, location, and confidence. The above technical solutions not only address interference from complex background textures but also improve the detection model's accuracy and robustness in identifying cross-material and multi-morphological defects, thereby enhancing the reliability of automated surface defect detection in industrial settings.
[0080] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0081] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0082] The above description is only a part of the embodiments of the present invention and does not limit the scope of protection of the present invention. Any equivalent device or equivalent process transformation made based on the content of the present invention specification and drawings, or direct or indirect application in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A direction-sensitive dual-branch industrial defect detection method, characterized in that, include: Acquire a surface image of the industrial product to be inspected, the surface image containing defect areas to be identified and background texture areas; The surface image is input into a preset orientation-sensitive dual-branch detection network, which includes a conventional semantic branch and an orientation texture branch. The surface image is feature extracted using the conventional semantic branch to generate a conventional semantic feature map. The surface image is feature extracted by the directional texture branch to generate a directional texture feature map. The directional texture branch uses the directional bias convolution operator to explicitly enhance and model the anisotropic texture patterns and linear defect structures in the surface image. The directional texture branch retains and aggregates the directional enhancement responses of multiple intermediate stages through the cascaded directional feature aggregation module to generate a continuously expressed directional texture feature map. The conventional semantic feature map and the directional texture feature map are fused at a preset key scale to obtain a multi-scale joint representation; The multi-scale joint representation is input into a multi-path cross-scale fusion neck network. The multi-path cross-scale fusion neck network performs cross-scale feature recombination on the multi-scale joint representation through a top-down semantic injection path and a bottom-up detail backflow path to generate enhanced multi-scale detection features. The enhanced multi-scale detection features are input into the detection head to generate and output the defect detection results of the industrial product to be inspected.
2. The direction-sensitive dual-branch industrial defect detection method according to claim 1, characterized in that, Feature extraction is performed on the surface image through the conventional semantic branch to generate a conventional semantic feature map, including: The surface image is input into the initial convolutional layer of the regular semantic branch, and the initial convolutional layer performs preliminary feature mapping on the surface image to generate initial semantic features. The initial semantic features are sequentially passed to multiple stacked semantic bottleneck modules. Each semantic bottleneck module contains several standard convolutional layers and residual connection structures. The standard convolutional layers use isotropic convolutional kernels to abstract and compress the initial semantic features. The residual connection structures add the input and output features of the semantic bottleneck modules element by element to gradually extract structural information with high-level semantic discriminative ability while maintaining stable feature transmission, thereby generating deep semantic features. During the feature transfer process of the multiple stacked semantic bottleneck modules, the deep semantic features of multiple preset intermediate stages are retained and used as multi-scale semantic candidate features. The multi-scale semantic candidate features are aggregated across levels with the final output features of the regular semantic branch to generate a regular semantic feature map containing rich contextual information and spatial details.
3. The direction-sensitive dual-branch industrial defect detection method according to claim 1, characterized in that, Feature extraction is performed on the surface image through the directional texture branch to generate a directional texture feature map, including: The surface image is input into the initial directional convolutional layer of the directional texture branch. The initial directional convolutional layer performs local directional structure perception on the surface image through the directional bias convolution operator to generate initial directional texture features. The initial directional texture features are input to a cascaded directional feature aggregation module. The cascaded directional feature aggregation module includes a main branch and a reserved branch. The main branch sequentially passes the initial directional texture features to multiple cascaded directional bottleneck modules. Each directional bottleneck module performs directional enhancement and updates the features through a directional bias convolution operator to generate multiple intermediate-stage directional enhancement features. The reserved branch performs an identity mapping on the initial directional texture features to generate reserved branch features. The preserved branch features, the initial directional texture features, and all the intermediate stage directional enhancement features are concatenated in the channel dimension to generate multi-stage directional aggregation features. The multi-stage directional aggregation features are reduced in dimensionality and integrated with information by channel compression convolution to generate a continuously expressive directional texture feature map.
4. The direction-sensitive dual-branch industrial defect detection method according to claim 3, characterized in that, The initial directional texture features are input into a cascaded directional feature aggregation module. This module includes a main branch and a retention branch. The main branch sequentially passes the initial directional texture features to multiple cascaded directional bottleneck modules. Each directional bottleneck module performs directional enhancement updates on the features using a directional bias convolution operator, generating multiple intermediate-stage directional enhancement features, including: The initial directional texture features are input into the main branch; In the main branch, the initial directional texture features are input to the first directional bottleneck module. The first directional bottleneck module includes a first directional bias convolutional layer and a second directional bias convolutional layer connected in series. The first directional bias convolutional layer uses an asymmetric convolutional kernel to enhance the main directional structural response of the input features, generating main directional enhanced features. The second directional bias convolutional layer uses an orthogonal directional asymmetric convolutional kernel to perform orthogonal directional context fusion on the main directional enhanced features, generating a first intermediate directional enhanced feature. The initial directional texture features and the first intermediate directional enhanced features are then added element-wise through a residual connection to output the first directional bottleneck module features. The first directional bottleneck module features are input into the second directional bottleneck module, which includes a third directional biased convolutional layer and a fourth directional biased convolutional layer in series. The third directional biased convolutional layer uses an asymmetric convolutional kernel to enhance the main directional structural response of the first directional bottleneck module features, generating a second main directional enhanced feature. The fourth directional biased convolutional layer uses an orthogonal directional asymmetric convolutional kernel to perform orthogonal directional context fusion on the second main directional enhanced feature, generating a second intermediate directional enhanced feature. The first directional bottleneck module features and the second intermediate directional enhanced feature are then added element-wise through a residual connection to output the second directional bottleneck module features. Repeat the above steps, passing the features output by the previous bottleneck module to each subsequent bottleneck module in sequence. Each bottleneck module performs orientation enhancement updates through the orientation bias convolution operator and retains the input features within the module through residual connections, until the feature transfer of all serial bottleneck modules is completed, generating orientation enhancement features for all intermediate stages. The orientation enhancement features for all intermediate stages include the intermediate orientation enhancement features output by each bottleneck module.
5. The direction-sensitive dual-branch industrial defect detection method according to claim 1, characterized in that, The conventional semantic feature map and the directional texture feature map are fused at a preset key scale to obtain a multi-scale joint representation, including: Obtain the conventional semantic feature map output by the conventional semantic branch at a preset key scale, and the directional texture feature map output by the directional texture branch at the same preset key scale; Determine whether the spatial resolution of the conventional semantic feature map and the directional texture feature map are consistent; If the spatial resolutions are consistent, the conventional semantic feature map and the directional texture feature map are directly concatenated in the channel dimension to generate an initial dual-branch fusion feature. If the spatial resolutions are inconsistent, an adaptive average pooling operation is used to downsample the higher-resolution feature map in the conventional semantic feature map and the directional texture feature map to the same target size as the lower-resolution feature map. Then, the two are concatenated in the channel dimension to generate the initial dual-branch fusion feature. The initial dual-branch fusion feature is subjected to channel dimensionality reduction and information recombination by a preset channel compression convolution to generate the dual-branch fusion feature at the preset key scale. Repeat the above steps to obtain the bi-branch fusion features corresponding to each of the multiple preset key scales, and use all the bi-branch fusion features together as a multi-scale joint representation.
6. The direction-sensitive dual-branch industrial defect detection method according to claim 5, characterized in that, The initial dual-branch fusion feature is subjected to channel dimensionality reduction and information recombination through a preset channel compression convolution to generate the dual-branch fusion feature at the preset key scale, including: The initial dual-branch fusion feature is obtained, which is formed by concatenating the conventional semantic feature map and the directional texture feature map in the channel dimension, and its number of channels is the sum of the number of channels of the conventional semantic feature map and the number of channels of the directional texture feature map; The initial dual-branch fusion feature is input into a preset channel compression convolutional layer. The channel compression convolutional layer uses a convolutional kernel of a preset size to perform feature interaction and information extraction in the local spatial neighborhood of the initial dual-branch fusion feature to generate intermediate compressed features. The intermediate compressed features are linearly transformed and non-linearly activated in the channel dimension by the channel compression convolution layer, thereby compressing the number of channels of the intermediate compressed features to a preset target number of channels and generating a compact fused feature after channel compression. The compact fusion features are batch normalized by the channel compression convolutional layer, and the numerical distribution of the compact fusion features is normalized and adjusted to generate normalized fusion features. The normalized fusion features are mapped using a nonlinear activation function through the channel compression convolutional layer, introducing nonlinear transformation capability to generate the dual-branch fusion features at the preset key scale.
7. The direction-sensitive dual-branch industrial defect detection method according to claim 1, characterized in that, The multi-scale joint representation is input into a multi-path cross-scale fusion neck network. This network performs cross-scale feature reorganization on the multi-scale joint representation through a top-down semantic injection path and a bottom-up detail backflow path, generating enhanced multi-scale detection features, including: The multi-scale joint representation is obtained, which includes high-resolution detail features, mid-level structural features, and high-level semantic features. By using a top-down semantic injection path, the high-level semantic features are upsampled so that the spatial resolution of the upsampled high-level semantic features is aligned with that of the mid-level structural features, thus obtaining the upsampled high-level semantic features. The upsampled high-level semantic features are concatenated with the mid-level structural features, and the concatenation result is locally reorganized using a convolutional block to generate enhanced mid-level features. The enhanced mid-layer features are upsampled so that the upsampled enhanced mid-layer features are aligned with the spatial resolution of the high-resolution detail features. The upsampled enhanced mid-layer features are concatenated with the high-resolution detail features, and the concatenation block is used to reorganize the local features of the concatenation result to generate enhanced shallow features. By using a bottom-up detail reflow path, the enhanced shallow features are downsampled so that the spatial resolution of the downsampled enhanced shallow features is aligned with that of the enhanced mid-layer features. The downsampled enhanced shallow features are concatenated with the enhanced mid-layer features, and the concatenation result is locally reorganized using a convolutional block to generate updated mid-layer features. The updated mid-level features are downsampled to align the spatial resolution of the downsampled updated mid-level features with that of the high-level semantic features. The downsampled updated mid-layer features are concatenated with the high-layer semantic features, and the concatenation result is locally reorganized through a convolutional block to generate updated high-layer features. The enhanced shallow features, the updated middle-layer features, and the updated high-layer features are used together as the enhanced multi-scale detection features.
8. The direction-sensitive dual-branch industrial defect detection method according to claim 7, characterized in that, The high-level semantic features are upsampled using a top-down semantic injection path, aligning the spatial resolution of the upsampled high-level semantic features with that of the mid-level structural features, resulting in upsampled high-level semantic features, including: The high-level semantic features and the mid-level structural features are obtained. The high-level semantic features have low spatial resolution and strong semantic abstraction information, while the mid-level structural features have high spatial resolution and richer structural detail information. Determine the spatial resolution ratio between the high-level semantic features and the mid-level structural features, wherein the spatial resolution ratio is a multiplier obtained by dividing the spatial size of the mid-level structural features by the spatial size of the high-level semantic features; Based on the spatial resolution ratio, a corresponding upsampling mode is selected, which includes one of nearest neighbor interpolation upsampling, bilinear interpolation upsampling, or transposed convolution upsampling. The high-level semantic features are spatially amplified using the upsampling mode, so that the spatial size of the amplified high-level semantic features is completely consistent with the spatial size of the mid-level structural features, thereby generating spatially aligned high-level semantic features. The spatially aligned high-level semantic features are output as upsampled high-level semantic features.
9. The direction-sensitive dual-branch industrial defect detection method according to claim 1, characterized in that, The enhanced multi-scale detection features are input into the detection head to generate and output the defect detection results of the industrial product to be inspected, including: The enhanced multi-scale detection features are obtained, which include enhanced shallow features, updated middle-level features, and updated high-level features; The enhanced shallow features, the updated middle-layer features, and the updated high-layer features are respectively input into the corresponding classification and regression branches in the detection head; The feature response at each preset anchor point is classified and predicted by the classification branch to generate the classification confidence score of each preset anchor point belonging to the preset defect category; The regression branch is used to predict the bounding box offset of the feature response at each preset anchor point position, thereby generating the predicted bounding box coordinates corresponding to each preset anchor point position. Based on the classification confidence, non-maximum suppression is applied to all the predicted bounding boxes to remove redundant predicted bounding boxes with high overlap and low confidence. The predicted bounding boxes retained after filtering, their corresponding defect categories, and classification confidence scores are correlated and integrated to generate and output the defect detection results of the industrial product to be inspected.
10. A direction-sensitive dual-branch industrial defect detection system, characterized in that, The system applicable to the method of any one of claims 1 to 9 comprises: An image acquisition module is used to acquire a surface image of an industrial product to be inspected, the surface image containing a defect area to be identified and a background texture area; A direction-sensitive dual-branch detection network module includes a regular semantic branch and a directional texture branch. The regular semantic branch is used to extract features from the surface image and generate a regular semantic feature map. The directional texture branch is used to extract features from the surface image and generate a directional texture feature map. The directional texture branch is internally configured with a directional bias convolution operator and a cascaded directional feature aggregation module. The directional bias convolution operator is used to explicitly enhance and model anisotropic texture patterns and linear defect structures in the surface image. The cascaded directional feature aggregation module is used to retain and aggregate the directional enhancement responses of multiple intermediate stages to generate a continuously expressed directional texture feature map. The key scale fusion module is used to fuse the conventional semantic feature map and the directional texture feature map at a preset key scale to generate a multi-scale joint representation. A multi-path cross-scale fusion neck network module is used to receive the multi-scale joint representation and perform cross-scale feature recombination on the multi-scale joint representation through a top-down semantic injection path and a bottom-up detail backflow path to generate enhanced multi-scale detection features. The detection head module is used to receive the enhanced multi-scale detection features, generate and output the defect detection results of the industrial product to be inspected.