Surface defect detection method and surface defect detection device

A surface defect detection method that combines adaptive weight parameters, dynamic convolution, and shared convolution with learnable scale parameters solves the problems of low detection accuracy and high missed detection rate in traditional detection methods, and achieves efficient detection of complex defect shapes.

CN120655644APending Publication Date: 2025-09-16HANGZHOU DIANZI UNIV +1

Patent Information

Application Number
CN202511150693.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Traditional surface defect detection methods have low detection accuracy, low efficiency, and poor adaptability, making it difficult to meet the needs of large-scale automated production in modern industry, especially for complex defect shapes, where the detection accuracy is low and the missed detection rate is high.

Method used

A surface defect detection method using adaptive weight parameters, dynamic convolution and shared convolution combined with learnable scale parameters is proposed to optimize feature expression and prediction results through feature extraction, fusion and prediction.

Benefits of technology

The accuracy of surface defect detection is improved, the missed detection rate is reduced, the adaptability to complex defect shapes is enhanced, and calculation redundancy is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655644A_ABST
    Figure CN120655644A_ABST
Patent Text Reader

Abstract

The invention relates to a surface defect detection method and a surface defect detection device. The method comprises the following steps: performing feature extraction and feature fusion on a to-be-detected surface image to obtain a fused feature; fusing the fusion features based on adaptive weight parameters to obtain joint features; based on dynamic convolution and shared convolution, respectively predicting the joint features to obtain a positioning prediction result and a classification prediction result; and adjusting the positioning prediction result and the classification prediction result based on learnable scale parameters to obtain a defect detection result. By adopting the method, the accuracy of surface defect detection can be improved, and the omission ratio can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of target detection technology, and in particular to a surface defect detection method and a surface defect detection device. Background Art

[0002] In the industrial field, surface defects of equipment (such as cracks, inclusions, and plaques) not only directly affect its service life but may also pose safety risks. Therefore, efficient and accurate surface defect detection is a key link in ensuring equipment quality.

[0003] Traditional technologies rely on manual visual inspection or optical imaging to detect defects on equipment surfaces. This often results in low accuracy, inefficiency, and poor adaptability, making it difficult to meet the demands of large-scale automated production in modern industry. While convolutional neural networks (CNNs) have proven effective at defect detection with the advancement of deep learning, the geometric diversity of surface defects, such as long, curved cracks and low-contrast linear scratches, makes CNNs inadequate for this diverse range of defects. This results in low accuracy and a high rate of missed detections. Summary of the Invention

[0004] Based on this, it is necessary to provide a surface defect detection method and a surface defect detection device that can improve the accuracy of surface defect detection and reduce the missed detection rate in order to address the above technical problems.

[0005] In a first aspect, the present application provides a surface defect detection method, the surface defect detection method comprising:

[0006] Perform feature extraction and feature fusion on the surface image to be detected to obtain fusion features;

[0007] fusing the fused features based on an adaptive weight parameter to obtain a joint feature;

[0008] Based on dynamic convolution and shared convolution, the joint features are predicted respectively to obtain positioning prediction results and classification prediction results;

[0009] Based on the learnable scale parameter, the positioning prediction result and the classification prediction result are adjusted to obtain a defect detection result.

[0010] In one embodiment, the extracting and fusing features of the surface image to be detected to obtain the fused features includes:

[0011] Performing feature extraction on the surface image to be detected based on a pre-trained backbone network to obtain image extraction features of multiple different scales;

[0012] Based on the pre-trained neck network, feature fusion is performed on the plurality of image extraction features to obtain the fused feature.

[0013] In one embodiment, the backbone network includes multiple multi-scale feature extraction modules, which are used to: perform dynamic convolution processing based on input features to obtain multi-scale convolution features; and perform splicing based on the multi-scale convolution features and the input features to obtain output features;

[0014] The multi-scale feature extraction module includes a variable convolution module, which is used to: determine the target geometric direction of the input feature based on the input feature and the persistent homology tool; determine the offset parameter based on multiple sampling points of the input feature and the target geometric direction; determine the matching convolution kernel and sampling path based on the offset parameter; and perform dynamic convolution processing on the input feature based on the convolution kernel and the sampling path to obtain the multi-scale convolution feature.

[0015] In one embodiment, determining the offset parameter based on the plurality of sampling points of the input feature and the target geometric orientation includes:

[0016] Based on each of the sampling points, determining an offset corresponding to each sampling point;

[0017] Based on the target geometric direction of the input feature, the offsets of the plurality of sampling points are accumulated to obtain the offset parameter.

[0018] In one embodiment, the neck network includes a feature fusion diffusion pyramid network for:

[0019] Performing multi-scale processing and parallel depth convolution based on the features extracted from the multiple images to obtain high-density semantic information;

[0020] Based on an adaptive weight parameter, the high-density semantic information is weightedly superimposed with the plurality of image extraction features to obtain a plurality of fusion features.

[0021] In one embodiment, the feature fusion diffusion pyramid network includes a feature focusing module, the feature focusing module includes a multi-scale processing module and a parallel depth convolution module, and the multi-scale processing module is used to:

[0022] Based on the scale of the features extracted from each image, feature processing is performed to match the scale to obtain a preset scale feature; the feature processing includes at least one of upsampling, downsampling, and convolution;

[0023] splicing a plurality of the preset scale features to obtain a focus feature;

[0024] The focused features are input into the parallel depthwise convolution module to obtain the high-density semantic information.

[0025] In one embodiment, the predicting of the joint features based on dynamic convolution and shared convolution to obtain positioning prediction results and classification prediction results includes:

[0026] Generating geometric adaptability parameters based on the joint features; adjusting the sampling position and weight of the convolution kernel based on the geometric adaptability parameters to obtain a positioning prediction result;

[0027] The classification weight is dynamically adjusted based on the response feature of the joint feature to obtain a classification prediction result.

[0028] In one embodiment, the geometric adaptability parameters include a spatial weight mask and a position offset parameter, and adjusting the sampling position and weight of the convolution kernel based on the geometric adaptability parameters to obtain a positioning prediction result includes:

[0029] performing weighted processing on the sampling points in the joint feature based on the spatial weight mask;

[0030] Adjusting the sampling position of the variable convolution kernel based on the position offset parameter;

[0031] The joint features are processed based on the adjusted convolution kernel to obtain a positioning prediction result.

[0032] In one embodiment, dynamically adjusting the classification weight based on the response feature of the joint feature to obtain the classification prediction result includes:

[0033] Dimensionally compressing the joint features based on a shared convolution structure to obtain compressed features;

[0034] Based on the response strength of the compressed features, a learnable response scaling layer is used to dynamically adjust the classification weights of multi-scale objects;

[0035] Based on the adjusted classification weights, a classification prediction result is determined; the classification prediction result includes a defect category and a confidence level corresponding to the defect category.

[0036] In a second aspect, the present application provides a surface defect detection device, comprising:

[0037] A preprocessing module is used to extract and fuse features of the surface image to be detected to obtain fused features;

[0038] A fusion module, configured to fuse the fusion features based on an adaptive weight parameter to obtain a joint feature;

[0039] A prediction module, configured to predict the joint features based on dynamic convolution and shared convolution, respectively, to obtain a positioning prediction result and a classification prediction result;

[0040] An adjustment module is used to adjust the positioning prediction result and the classification prediction result based on a learnable scale parameter to obtain a defect detection result.

[0041] The above-mentioned surface defect detection method and surface defect detection device obtain fused features by extracting and fusing features of the surface image to be detected; fuse the fused features based on adaptive weight parameters to obtain joint features; predict the joint features based on dynamic convolution and shared convolution respectively to obtain positioning prediction results and classification prediction results; adjust the positioning prediction results and the classification prediction results based on learnable scale parameters to obtain defect detection results. Multi-scale information can be retained through multi-layer convolution feature extraction and fusion; dynamic weighted optimization of features can be achieved based on fusion of adaptive weight parameters; dynamic convolution and shared convolution can be used to improve adaptability to complex defect shapes and reduce computational redundancy. Adjusting the prediction results through learnable parameters can compensate for model deviations, thereby achieving the technical effect of improving the accuracy of surface defect detection and reducing the missed detection rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 A diagram showing an application environment of a surface defect detection method in one embodiment;

[0043] Figure 2 Schematic diagram of a process for detecting surface defects in one embodiment;

[0044] Figure 3 is a structural block diagram of a defect detection model in one embodiment;

[0045] Figure 4 is a structural block diagram of a C3k2 module in one embodiment;

[0046] Figure 5 A diagram of a dynamic serpentine convolution deformation process in one embodiment;

[0047] Figure 6 This is a diagram showing the result of dynamic snake-shaped convolution deformation in one embodiment;

[0048] Figure 7 is a structural block diagram of a feature focusing module in one embodiment;

[0049] Figure 8 is a schematic diagram of the diffusion of focusing features in one embodiment;

[0050] Figure 9 is a structural block diagram of a dynamic task alignment detection head in one embodiment;

[0051] Figure 10 2 is a block diagram of a convolution group normalization module according to an embodiment;

[0052] Figure 11 is a structural block diagram of a task decomposition module in one embodiment;

[0053] Figure 12 is a structural block diagram of a surface defect detection device in one embodiment;

[0054] Figure 13 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0055] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0056] The surface defect detection method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store data that server 104 needs to process. The data storage system can be integrated with server 104, or placed on a cloud or other network server. By communicating with server 104, terminal 102 can obtain an image of the surface to be inspected, perform feature extraction and feature fusion on the image to be inspected, and obtain fused features. The fused features are then fused based on adaptive weight parameters to obtain joint features. The joint features are then predicted based on dynamic convolution and shared convolution to obtain positioning prediction results and classification prediction results. The positioning prediction results and classification prediction results are then adjusted based on a learnable scale parameter to obtain defect detection results. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart car devices, etc. Portable wearable devices can include smart watches, smart bracelets, head-mounted devices, etc. Server 104 can be implemented as a standalone server or a server cluster consisting of multiple servers.

[0057] In one embodiment, Figure 2 As shown, a surface defect detection method is provided. Taking the method applied to the terminal 102 as an example, the surface defect detection method includes the following steps:

[0058] Step S100 : performing feature extraction and feature fusion on the surface image to be detected to obtain fused features.

[0059] Feature extraction can be the extraction of discriminative visual information from the original image through a specific algorithm, which may include but is not limited to texture, edge, shape, etc.

[0060] Feature extraction can be achieved through convolutional neural networks (such as ResNet, VGG) or multi-layer convolution operations. In an exemplary embodiment, edges and textures can be captured by the shallow layer of the convolutional neural network, and complex shapes or defect patterns can be captured by the deep layer.

[0061] Feature fusion can be the integration of feature information from different sources or at different levels to enhance expressiveness. For example, feature fusion can include one or more processing processes such as weighted summation, concatenation, and attention mechanism.

[0062] The fused feature may be a multi-scale feature set after feature fusion, and may include, for example, a high-dimensional feature representation merged by channel splicing.

[0063] In an exemplary embodiment, low-level to high-level features may be extracted layer by layer through a multi-layer convolutional network, and then the importance weights of each feature layer may be dynamically calculated and weighted summed using an attention mechanism, or multi-scale features may be merged through channel splicing, thereby obtaining a fusion feature with less redundancy and retaining multi-scale information.

[0064] Step S200: Fusing the fused features based on the adaptive weight parameters to obtain joint features.

[0065] The adaptive weight parameters may be parameters that dynamically adjust the weights of different features through learnable parameters. For example, the adaptive weight parameters may include input-related weight coefficients generated by a fully connected layer or a preceding neural network.

[0066] The joint features may be optimized features that have been dynamically weighted and recombined. For example, the joint features may include a weighted feature set obtained by assigning a higher weight to key area features.

[0067] In an exemplary embodiment, the joint features can be obtained by generating weight coefficients through a lightweight network, multiplying the weight coefficients with the original fusion features to achieve dynamic weighting, and then recombining the weighted features to enhance the key area representation capability and reduce irrelevant information interference.

[0068] In step S300 , the joint features are predicted based on dynamic convolution and shared convolution to obtain positioning prediction results and classification prediction results.

[0069] Dynamic convolution can be a convolution process in which the parameters of the convolution kernel are dynamically adjusted according to the input features. Exemplarily, dynamic convolution can adopt one or more of deformable convolution and conditional convolution.

[0070] Shared convolution can be a convolution layer that uses the same set of convolution parameters for multiple tasks. For example, it can include the bottom convolution layer shared by the localization branch and the classification branch.

[0071] The positioning prediction result may be prediction information of the defect location, and may include bounding box coordinates, for example. The classification prediction result may be prediction information of the defect type, and may include category probability distribution, for example.

[0072] In an exemplary embodiment, positioning prediction results and classification prediction results can be obtained by adjusting the sampling point offset through deformable convolution to capture complex defect shapes, or generating convolution kernels related to the defect type through conditional convolution, thereby achieving positioning prediction; after extracting common features through a shared convolution layer, the common features are input into the positioning sub-network and the classification sub-network respectively, thereby taking into account both adaptability and computational efficiency.

[0073] Step S400 : Based on the learnable scale parameter, the positioning prediction result and the classification prediction result are adjusted to obtain a defect detection result.

[0074] The learnable scale parameter can be a parameter learned through backpropagation and used to adjust the scale or ratio of the prediction result. For example, the learnable scale parameter can include a learnable scaling factor or offset. The defect detection result can be the final integrated defect location, type, and confidence information.

[0075] In an exemplary embodiment, defect detection results may be obtained by correcting the deviation of bounding box coordinates through learnable scaling factors and offsets, and by calibrating classification probabilities through temperature scaling, thereby compensating for model deviations and reducing missed or false detections.

[0076] This embodiment provides a surface defect detection method, which obtains fused features by performing feature extraction and feature fusion on a surface image to be detected; obtains joint features by fusing the fused features based on adaptive weight parameters; predicts the joint features based on dynamic convolution and shared convolution to obtain positioning prediction results and classification prediction results; adjusts the positioning prediction results and the classification prediction results based on learnable scale parameters to obtain defect detection results. Multi-scale information can be retained through multi-layer convolution feature extraction and fusion; dynamic weighted optimization of features can be achieved through fusion based on adaptive weight parameters; dynamic convolution and shared convolution can be used to improve adaptability to complex defect shapes and reduce computational redundancy. Adjusting the prediction results through learnable parameters can compensate for model deviations, thereby achieving the technical effect of improving the accuracy of surface defect detection and reducing the missed detection rate.

[0077] In one embodiment, feature extraction and feature fusion are performed on the surface image to be detected, and the fused features obtained include:

[0078] Based on the pre-trained backbone network, feature extraction is performed on the surface image to be detected, and image extraction features of multiple different scales are obtained;

[0079] Based on the pre-trained neck network, feature fusion is performed on multiple image extraction features to obtain fused features.

[0080] The pre-trained backbone network can be a convolutional neural network pre-trained on a labeled dataset, used to extract hierarchical features of the image. Furthermore, the backbone network can be obtained through transfer learning or model loading. Exemplarily, the pre-trained backbone network can include one or more of ResNet, VGG, or EfficientNet. Through multiple layers of convolution, the backbone network can abstract local to global features of the image layer by layer, for example, extracting edge textures at shallow layers and complex shapes or defect patterns at deeper layers.

[0081] In an exemplary embodiment, feature extraction is performed based on a pre-trained backbone network, and an image can be input into the backbone network. The backbone network can extract low-level local features of the input image through an initial convolutional layer, and gradually reduce the spatial resolution of the feature map and increase the number of channels through a downsampling module, thereby outputting multiple feature maps with different scales at different levels, such as 1 / 4, 1 / 8, 1 / 16 of the input image size, etc.

[0082] The pre-trained neck network can be used to integrate feature maps from different levels to achieve multi-scale feature fusion. Exemplarily, the pre-trained neck network can include one or more of a feature pyramid network, a path aggregation network, or a bidirectional feature pyramid network. The neck network generates multi-scale fused features through one or more operations including upsampling, downsampling, and concatenation, thereby enhancing the model's ability to detect defects of varying sizes.

[0083] In one exemplary embodiment, feature fusion based on a pre-trained neck network can be performed through a top-down pathway, where high-level features are upsampled and concatenated with low-level features to supplement detailed information. For example, deconvolution or interpolation can be performed through a feature pyramid network, where high-level feature maps are upsampled and concatenated with low-level feature maps. Lateral connections are used to align the channel dimensions, and features are further fused through convolutional layers. In other examples, a bidirectional pathway can be used in conjunction with bottom-up information transfer to output a fused feature map that contains both high-resolution details and deep semantic information.

[0084] This embodiment provides a surface defect detection method, which extracts multi-scale features based on a pre-trained backbone network to cover defect patterns from local texture to global structure, combines the pre-trained neck network to fuse features at different levels to enhance the complementarity of semantics and details, and optimizes computational efficiency through lightweight design, and utilizes pre-trained weights to improve the generalization ability of the model. The model can adapt to the geometric diversity of defects, and feature fusion can also enhance the robustness of detection under changes in illumination or scale, thereby achieving the technical effect of improving the accuracy of surface defect detection and reducing the missed detection rate.

[0085] In one embodiment, the backbone network includes multiple multi-scale feature extraction modules, which are used to: perform dynamic convolution processing based on input features to obtain multi-scale convolution features; and concatenate the multi-scale convolution features and input features to obtain output features;

[0086] The multi-scale feature extraction module includes a variable convolution module, which is used to: determine the target geometric direction of the input feature based on the input feature and the persistent homology tool; determine the offset parameter based on multiple sampling points of the input feature and the target geometric direction; determine the matching convolution kernel and sampling path based on the offset parameter; and perform dynamic convolution processing on the input feature based on the convolution kernel and the sampling path to obtain multi-scale convolution features.

[0087] Among them, the multi-scale feature extraction module is used to generate multi-scale feature representation through dynamic convolution and feature splicing. Exemplarily, the multi-scale feature extraction module may include a variable convolution module. Further, the multi-scale feature extraction module may also include a feature splicing module.

[0088] The input feature may be a feature map output by a preceding network layer. For example, the input feature may be generated by one or more network layers such as convolution operations and nonlinear activation functions.

[0089] Dynamic convolution processing can be a convolution operation method that performs convolution after adaptively adjusting the parameters of the convolution kernel and the sampling path. Exemplarily, dynamic convolution processing includes offset parameter calculation, convolution kernel and path determination, and convolution operation.

[0090] Multi-scale convolutional features can be feature representations that fuse information from different spatial scales. For example, multi-scale convolutional features can include mixed features of local texture and global contour. Concatenation can be a method of fusing features along the channel dimension. For example, this can be achieved by concatenating the original input features with the processed features.

[0091] The output feature may be a comprehensive representation including original details and geometric adaptive features. In an exemplary embodiment, the number of channels of the output feature may be the sum of the number of channels of the input feature and the multi-scale convolution feature.

[0092] In one exemplary embodiment, when input features are passed to the backbone network, they may include a process of being passed to a variable convolution module, where the variable convolution module is used to analyze the local geometric orientation of the input features and the results of the persistent homology analysis, thereby calculating the offset parameters to adjust the spatial distribution of the sampling points; then, the offset parameters are combined with the topological importance assessed by the persistent homology tool to generate convolution kernel parameters and sampling paths that match the offset parameters; dynamic convolution operations are performed based on the determined convolution kernel parameters and sampling paths to obtain multi-scale convolution features. The multi-scale convolution features can then be spliced ​​with the original input features along the channel dimension to form output features that contain the original details and geometric adaptability features, thereby achieving the technical effect of enhancing the integrity of feature expression and geometric adaptability. In particular, for the detection of bending cracks, sampling can be performed along the defect orientation while retaining high-resolution details, thereby improving positioning accuracy.

[0093] Furthermore, the variable convolution module may be a submodule for dynamically adjusting convolution parameters, and may be embedded in the hierarchy of the multi-scale feature extraction module.

[0094] The multiple sampling points of the input features may be a set of discrete spatial positions on the feature map. For example, the distribution of the sampling points may be a regular grid or a non-uniform distribution after adaptive offset.

[0095] The persistent homology tool can be a mathematical method for quantifying the stability of topological features. For example, the persistent homology tool can be used to calculate the persistence of connected regions or holes at different scales. The target geometric orientation can be the spatial distribution direction of the defect structure. For example, the target geometric orientation can be derived from the results of the persistent homology analysis.

[0096] The offset parameter may be a set of numerical values ​​that describe the displacement of the sampling point. For example, the offset parameter may include displacement components in the horizontal and vertical directions.

[0097] The convolution kernel may be an adaptively generated weight matrix. For example, the shape of the convolution kernel may be related to the spatial distribution of the sampling path. The sampling path may be a spatial trajectory of the convolution operation. For example, the sampling path may be an irregular sampling trajectory formed along the defect direction.

[0098] In an exemplary embodiment, offset parameters describing the displacement of sampling points can be calculated based on multiple sampling points of the input feature and their local gradient directions, such as guiding the sampling points to offset along the curvature direction at the crack bend; the topological structure of the input feature can be analyzed in combination with the persistent homology tool to identify geometric patterns with high persistence to determine the target geometric direction and assign priority sampling paths, such as generating dense sampling trajectories for continuous crack path areas; dynamic convolution operations can be performed based on the generated convolution kernel weights and the sampling path, such as extracting multi-scale features along the optimized path, thereby achieving the technical effect of improving topological perception and geometric matching, ignoring short-term noise areas and strengthening the feature expression of key defect structures, while reducing the computational overhead of fully adaptive convolution.

[0099] This embodiment provides a surface defect detection method, which combines a multi-scale feature extraction module with a variable convolution module and a feature stitching operation, uses a persistent homology tool to analyze the topological structure to optimize the convolution parameters, and adjusts the sampling path through offset parameters to match the target geometric direction, and fuses the dynamic convolution features with the original input features. This can achieve the technical effects of enhancing the adaptability of complex defect geometry, improving the integrity of multi-scale feature expression, reducing the missed detection rate, and improving detection accuracy while maintaining computational efficiency.

[0100] In one embodiment, determining the offset parameter based on a plurality of sampling points of the input feature and the target geometric orientation includes:

[0101] Based on each sampling point, determining an offset corresponding to each sampling point;

[0102] Based on the target geometric direction of the input feature, the offsets of multiple sampling points are accumulated to obtain the offset parameters.

[0103] The offset can describe the displacement of a single sampling point relative to its original grid position. For example, the offset can be used to adjust the convolution kernel's sampling path to match the local geometric features of the defect. The offset can be determined by prediction using a lightweight neural network, such as a small fully connected layer or convolutional layer. Furthermore, the offset can include two-dimensional displacement values, such as lateral and longitudinal offsets.

[0104] The multiple sampling points of the input features can be pixels on a grid or feature map locations. The target geometric orientation can be the overall geometric trend or direction of the defect, such as the curved path of a crack or the extension direction of a scratch, which can be derived from topological features extracted by persistent homology tools. Furthermore, the target geometric orientation can include the main direction of a connected region or the statistical results of the direction field.

[0105] Determining the offset based on each sampling point can be achieved through the analysis of local geometric features and the coordinated prediction of the offset. In a specific embodiment, geometric information can be extracted from the local area of ​​each sampling point, for example, the gradient direction, curvature or texture direction can be calculated to identify the local direction of the defect; and the offset of the point is predicted through a learnable neural network module, such as a small convolution layer. For example, if the local gradient direction indicates that the crack extends to the upper right, the predicted offset may include a displacement vector in the corresponding direction. It is understandable that the process of determining the offset also meets physical rationality constraints, such as limiting the maximum displacement amplitude to no more than 1 / 2 of the convolution kernel radius to avoid the sampling point exceeding the valid area.

[0106] The offsets are accumulated based on the target geometric direction, and the offsets of multiple sampling points can be accumulated through the target geometric direction. In a specific embodiment, the overall geometric direction of the input feature, such as the main path direction of the crack, can be determined by a persistent homology tool; the local offsets of all sampling points are arranged in spatial order, and cumulatively adjusted in combination with the geometric direction constraint. For example, if there is a conflict between the local offset directions of two adjacent points, they can be adjusted by weighted averaging or path smoothing algorithms to ensure the consistency of the overall direction; for key areas, such as crack inflection points, the local offset weights can be strengthened to improve the continuity of the path. Furthermore, the optimized offset sequence can be combined into a global offset parameter. Exemplarily, the offset parameter can be stored in the form of a matrix containing a displacement vector.

[0107] This embodiment provides a surface defect detection method that predicts the offset of each sampling point based on local geometric features to capture subtle changes in defects, and cumulatively optimizes the offset in combination with target geometric direction constraints to ensure path continuity. This method improves computational efficiency through staged processing, and through the collaborative analysis of local and global geometric features, enhances adaptability to slender curved defects such as serpentine cracks while reducing the risk of path breakage caused by local noise, thereby reducing the missed detection rate and improving the spatial sampling accuracy of dynamic convolution.

[0108] In one embodiment, the neck network includes a feature fusion diffusion pyramid network for:

[0109] Based on multiple image extraction features, multi-scale processing and parallel deep convolution are performed to obtain high-density semantic information;

[0110] Based on adaptive weight parameters, high-density semantic information is weightedly superimposed with multiple image extraction features to obtain multiple fusion features.

[0111] In this embodiment, the feature fusion diffusion pyramid network may be a feature pyramid network that enhances cross-level feature interaction through a diffusion mechanism.

[0112] Multi-scale processing can be an operation of spatial dimension alignment of feature maps of different resolutions. Exemplarily, it can be achieved through operations such as upsampling, downsampling, and channel unification. Exemplarily, it can also include operations such as bilinear interpolation and deconvolution.

[0113] Parallel depthwise convolution can be a convolution operation performed independently in multiple branches. In an exemplary embodiment, it can be implemented through convolution processes such as 3×3 convolution and deformable convolution. For example, the aligned feature maps can be used for texture extraction or structural enhancement respectively through convolution.

[0114] High-density semantic information can be a feature representation that includes enhanced semantic representation after deep processing. For example, it can be generated through a multi-branch convolution operation. Furthermore, high-density semantic information can include features that enhance defect category features or refine edge details.

[0115] The adaptive weight parameter can be a coefficient used to dynamically adjust the feature contribution. For example, the adaptive weight parameter can be obtained by a fully connected layer or an attention mechanism. Furthermore, the adaptive weight parameter can include a channel attention weight or a spatial weight matrix.

[0116] Weighted superposition can be an execution process of adding or concatenating feature maps element by element. For example, it can be a fusion of weight normalization and original features, such as weighted addition after Softmax normalization.

[0117] Multi-scale processing and parallel deep convolution based on feature extraction from multiple images can align the resolution of each feature map output by the backbone network. In each independent branch, deep convolution is used to extract high-density semantic information, thereby improving the depth and efficiency of feature expression. For example, deep features can enhance semantic category information through convolution, and shallow features can refine local texture details.

[0118] Based on adaptive weight parameters, high-density semantic information is weightedly superimposed with multiple image extraction features. The weight parameters can be calculated through a lightweight network. For example, the channel attention weight can be calculated for each level feature, and then the weighted high-density semantic information is fused with the original features, so as to dynamically optimize feature complementarity. It can be understood that when detecting low-contrast defects, the neck network of this embodiment can enhance the weight of shallow details, and when detecting large-area patches, the deep semantic weight is enhanced.

[0119] This embodiment provides a surface defect detection method that achieves spatial alignment and in-depth processing of multi-scale features through a feature fusion diffusion pyramid network. It dynamically adjusts feature fusion in combination with adaptive weight parameters. The feature fusion diffusion pyramid network enhances cross-level information interaction through parallel paths, and multi-scale processing and parallel convolution ensure that features at each level are fully optimized. Adaptive weight superposition dynamically enhances the contribution of key features based on the input content, thereby improving the multi-scale feature expression capability and fusion efficiency, and optimizing the positioning accuracy and classification reliability of defect detection.

[0120] In one embodiment, the feature fusion diffusion pyramid network includes a feature focusing module, which includes a multi-scale processing module and a parallel depth convolution module. The multi-scale processing module is used to:

[0121] Based on the scale of the features extracted from each image, feature processing is performed to match the scale to obtain a preset scale feature; the feature processing includes at least one of upsampling, downsampling, and convolution;

[0122] Multiple preset scale features are spliced ​​together to obtain a focused feature;

[0123] The focused features are input into the parallel depthwise convolution module to obtain high-density semantic information.

[0124] Among them, the multi-scale processing module can be a component responsible for adjusting the spatial resolution and optimizing the channels of image features at different levels. By extracting the scale of the image features, upsampling, downsampling or convolution operations can be dynamically selected to achieve the corresponding feature processing operations.

[0125] Exemplarily, upsampling can improve the spatial resolution of low-resolution features through sampling methods such as bilinear interpolation and transposed convolution. In an exemplary embodiment, upsampling can expand deep features from 1 / 16 of the input size to 1 / 4. Downsampling can reduce the spatial resolution of high-resolution features through sampling methods such as maximum pooling and strided convolution. In an exemplary embodiment, downsampling can compress shallow features from 1 / 4 to 1 / 8. The convolution operation can be a convolution operation such as 1×1 convolution, depthwise convolution, etc. to compress or expand the channel dimension. In an exemplary embodiment, the convolution operation of this embodiment can expand the number of channels from 256 to 512 to enhance feature expression capabilities.

[0126] The preset scale features can be the feature results of the image extraction features after the above processing, unified to a specific spatial size or number of channels. In one exemplary embodiment, through the above operation, the features at different levels can be adjusted to feature maps aligned with a preset resolution, such as 1 / 4 of the original image.

[0127] Feature stitching involves concatenating multiple pre-set scale features along the channel dimension. For example, assuming the input image extraction features include P3, P4, and P5 features, then after adjustment, these features have the same 1 / 4 spatial resolution. Their number of channels can be C2, C3, and C4, respectively, and the number of channels of the stitched focused feature can be C2+C3+C4. Feature stitching preserves the spatial alignment and information integrity of multi-scale features, integrating the semantic information of deep features with the detailed information of shallow features into a unified representation.

[0128] The parallel depth convolution module can be a component that includes multiple parallel convolution branches. For example, each branch can use a different convolution type to achieve differentiated feature extraction. In a specific embodiment, the first branch can use standard 3×3 convolution to capture local texture details, the second branch uses deformable convolution to adapt to irregular defect morphology, and the third branch uses depth-separable convolution to reduce computational complexity while extracting long-range semantic information. In addition, the parallel depth convolution module can also use other branch convolution processing methods, which is not limited to three branches, nor is it limited to the processing process of each branch mentioned above. This embodiment does not limit this. The results of the output of each branch can generate high-density semantic information through channel splicing or weighted fusion. In a specific embodiment, the output channels D1, D2, and D3 of the three branches can be merged into a feature representation of D1+D2+D3.

[0129] This embodiment provides a surface defect detection method that dynamically adjusts the resolution and optimizes channels of features at different levels through a multi-scale processing module, integrating scattered multi-scale information into a unified focused feature. High-density semantic information is then extracted through the multi-path convolution branch of a parallel deep convolution module, combined with differentiated operations such as standard convolution and deformable convolution. Targeted scale matching operations reduce redundant computations, while channel splicing preserves multi-scale complementarity. Parallel computing improves processing efficiency and feature density, effectively resolving semantic misalignment issues caused by resolution mismatch in traditional pyramid networks and the lack of adaptability of a single convolution type to complex defects. This enables precise positioning and classification of complex defects such as low-contrast patches and elongated cracks in industrial scenarios, reducing both missed detection and false detection rates.

[0130] In one embodiment, based on dynamic convolution and shared convolution, joint features are predicted respectively, and the positioning prediction results and classification prediction results obtained include:

[0131] Generate geometric adaptability parameters based on joint features;

[0132] Based on the geometric adaptability parameters, the sampling position and weight of the convolution kernel are adjusted to obtain the positioning prediction result;

[0133] The classification weight is dynamically adjusted based on the response characteristics of the joint features to obtain the classification prediction results.

[0134] The geometric adaptability parameters may be learnable parameters used to guide spatial sampling and weight adjustment of convolution kernels. For example, the geometric adaptability parameters may include a quantitative representation of geometric information such as defect edge direction and curvature. In some specific embodiments, the geometric adaptability parameters may include offset parameters in deformable convolution and / or spatial weight distributions generated by an attention mechanism. Furthermore, the geometric adaptability parameters may be extracted from joint features via a fully connected layer or a lightweight convolutional network.

[0135] The response feature of a joint feature can be a set of subfeatures within the joint feature that are highly responsive to a specific classification task. For example, the response feature can include feature channels that are highly correlated with the defect category. Furthermore, the response feature can be filtered through global pooling or a channel-wise attention mechanism.

[0136] Based on the geometric adaptability parameters, the sampling position and weight of the convolution kernel are adjusted to obtain the positioning prediction results. For example, the geometric adaptability parameters can be used as offsets through deformable convolution to make the convolution kernel sampling points deviate from the predefined grid and perform spatial sampling along the direction or edge of the defect; and the convolution kernel response is weighted by the spatial weight map generated by the geometric adaptability parameters to enhance the key area features, thereby improving the positioning accuracy of slender curved defects, such as realizing adaptive sampling of crack end points, curvature change areas, etc.

[0137] Dynamically adjusting the classification weight based on the response characteristics of the joint features can be done by calculating the weight of each feature channel through the channel attention mechanism to screen out sub-features that respond strongly to specific defect categories. Furthermore, the classification weight can be dynamically adjusted by similarity matching the sub-features with the preset category prototypes to strengthen the influence of key features, thereby reducing redundant feature interference. For example, in the classification of low-contrast defects, the weight of features related to light scratches can be amplified.

[0138] A surface defect detection method provided in this embodiment can explicitly model the geometric information of the defect by generating geometric adaptability parameters based on joint features, and use the geometric adaptability parameters to adjust the sampling position and weight of the convolution kernel to achieve adaptive positioning prediction. At the same time, the classification weights are dynamically optimized based on the response features to improve the category differentiation, thereby achieving the technical effect of enhancing the robustness of complex shape defect positioning and reducing interference between similar features between categories.

[0139] In one embodiment, the geometric adaptability parameters include a spatial weight mask and a position offset parameter. Based on the geometric adaptability parameters, the sampling position and weight of the convolution kernel are adjusted to obtain a positioning prediction result including:

[0140] Weighting the sampling points in the joint features based on the spatial weight mask;

[0141] Adjust the sampling position of the variable convolution kernel based on the position offset parameter;

[0142] The joint features are processed based on the adjusted convolution kernel to obtain the positioning prediction result.

[0143] The spatial weight mask can be a two-dimensional matrix that matches the spatial dimensions of the feature map. For example, each element can represent the feature response weight at the corresponding location. Furthermore, the spatial weight mask can be generated by modules or networks such as attention mechanisms, convolutional layers, or fully connected layers. For example, the spatial weight mask can be generated by combining global average pooling with deconvolution operations, or by filtering out the weights of defect-sensitive areas through a gating mechanism.

[0144] The position offset parameters can be a set of parameters that record the offset of the convolution kernel sampling points relative to the original grid. In one exemplary embodiment, they can be obtained by the offset prediction branch in the deformable convolution. By predicting the offset of each sampling point in the horizontal and vertical directions, an offset field can be formed.

[0145] The sampling points in the joint features are weighted based on a spatial weight mask. For example, a lightweight network module, such as a combination of a convolutional layer and a ReLU activation layer, can be used to generate a spatial weight mask. The spatial weight mask is then multiplied element-by-element by the joint features to achieve weighted processing. It is understandable that when the weight value of a region is high, the feature response of that region may be amplified; otherwise, it will be suppressed. By explicitly strengthening the feature response of the key defect area while weakening the interference of the background or noise area, the sensitivity of subsequent convolution operations to the defect can be improved.

[0146] The sampling position of the variable convolution kernel is adjusted based on the position offset parameter. For example, the position offset parameter can be generated by the offset prediction branch of the deformable convolution, and the sampling point coordinates of the original convolution kernel are added to the offset parameter to obtain the adjusted sampling position. It can be understood that if the defect edge is curved, the offset parameter will guide the sampling point to offset along the edge direction instead of being fixed on the grid node. The eigenvalues ​​of the sampling points can be obtained from the feature map by bilinear interpolation or other methods. By dynamically adjusting the sampling position, the convolution kernel can capture the geometric details of the defect more accurately and avoid positioning deviations caused by fixed sampling grids.

[0147] The joint features are processed based on the adjusted convolution kernel to obtain the positioning prediction result. For example, the convolution kernel that has been spatially weighted and position adjusted can be convolved with the feature map to generate an intermediate feature map for positioning prediction. In a specific embodiment, it can be obtained by outputting the bounding box coordinates of the regression layer, predicting the position of the defect center point through the heat map, etc. Then, by combining the dual adjustment results of the spatial weight and the position offset, the final positioning prediction can be made. The positioning prediction result acquisition process of this embodiment can enable the convolution kernel to have the ability to perceive key areas and adapt to geometric shapes at the same time. For example, for slender and curved cracks, the spatial weight can highlight the edge response, and the position offset ensures that the sampling points are distributed along the crack direction, thereby generating more accurate positioning results.

[0148] This embodiment provides a surface defect detection method that enhances the response of key areas by weighting joint features based on a spatial weight mask, dynamically adjusts the sampling position to adapt to the geometric shape based on a position offset parameter, and combines the double-adjusted convolution kernel to generate the final positioning prediction. This can improve the positioning accuracy and computational efficiency of complex-shaped defects and reduce the problem of mispositioning.

[0149] In one embodiment, dynamically adjusting the classification weight based on the response feature of the joint feature to obtain the classification prediction result includes:

[0150] Based on the shared convolution structure, the joint features are dimensionally compressed to obtain compressed features;

[0151] Based on the response strength of compressed features, a learnable response scaling layer is used to dynamically adjust the classification weights of multi-scale targets;

[0152] Based on the adjusted classification weights, a classification prediction result is determined; the classification prediction result includes a defect category and a confidence level corresponding to the defect category.

[0153] The shared convolutional structure can be a module that achieves feature dimension compression through lightweight convolution. Furthermore, the parameters of the shared convolutional structure can be obtained by reusing the convolutional layer parameters in the classification branch. Exemplarily, the shared convolutional structure can also include convolution kernels with learnable weights to reduce the number of feature channels or spatial resolution.

[0154] Dimensionality reduction can be achieved by reducing the number of channels or spatial dimensions of a feature map. Compressed features can be feature representations that retain key response information after dimensionality reduction. For example, compressed features can include local characteristic patterns of defect texture or shape.

[0155] Based on the response strength of compressed features, a learnable response scaling layer is used to dynamically adjust the classification weights of multi-scale targets. Response strength can be a quantitative indicator of the sensitivity of a feature region or channel to the classification task, and can be obtained through activation function output or attention mechanism. For example, the response strength can include the difference in activation values ​​between the defect center and the background region.

[0156] A learnable response scaling layer can be an adaptive structure that adjusts feature weights through a parameterized module. For example, the learnable response scaling layer can generate a scaling coefficient matrix that matches the response distribution. Multi-scale targets can be defect categories of different sizes in an image, such as patches and scratches of different sizes.

[0157] Dynamically adjusting classification weights, in a specific embodiment, can involve calculating the response strength distribution of compressed features, such as through global average pooling to obtain channel-level responses or through a spatial attention mechanism to obtain pixel-level responses; generating a scaling factor by inputting the response strength into a learnable scaling layer; and multiplying the scaling factor by the original classification weight to enhance the weight of high-response areas and suppress the weight of low-response areas. It can be understood that when a channel has a strong response to a certain category, the corresponding scaling factor can be increased, giving that channel a higher priority during classification.

[0158] A classification prediction result is determined based on the adjusted classification weight, wherein the adjusted classification weight may be a feature weight distribution result after dynamic scaling.

[0159] The classification prediction result can be a probability distribution containing defect categories and their confidence levels. For example, this can be obtained through the output of a fully connected layer or a classification head. In one specific embodiment, the classification prediction result can include a distribution with a probability of 0.8 for "crack" and a probability of 0.15 for "plaque." The confidence level can be a probability value that indicates the model's confidence in the prediction result and can be obtained by normalizing the classification probability.

[0160] This embodiment provides a surface defect detection method that achieves dimensionality compression through a shared convolutional structure, thereby reducing computational complexity and retaining key defect features. A learnable response scaling layer dynamically adjusts classification weights according to response strength, thereby enhancing the classification accuracy of multi-scale targets. This achieves the technical effects of improving the robustness of multi-scale defect classification, reducing computing resource consumption, suppressing background noise interference, and reducing false detection and missed detection rates.

[0161] In order to more clearly illustrate the technical solution of this application, this application also provides a detailed embodiment.

[0162] In one embodiment, a surface defect detection method is provided, comprising:

[0163] The surface image to be detected is input into the pre-trained defect detection model to obtain the defect detection result. In a specific embodiment, the defect detection model is trained based on the improved YOLO model. Figure 3 As shown, the defect detection model includes a backbone network, a neck network, and a detection head network.

[0164] The backbone network efficiently extracts multi-dimensional features from the input image, laying the foundation for subsequent detection tasks. Shallow convolutional layers capture low-level features such as edges and textures. Deeper layers further refine high-level features containing semantic information about the target. Image dimensionality is reduced through convolution and downsampling operations, generating feature maps of varying scales to meet the multi-scale object recognition requirements of object detection. These feature maps not only reduce computational complexity but also provide a rich feature representation for the neck network and detection head, supporting subsequent feature fusion, object localization, and classification.

[0165] The backbone network includes a C3K2 module that is improved based on the dynamic snake convolution module. Dynamic snake convolution is an innovative convolution method designed for feature extraction of slender, curved or tubular structures. At its core, it dynamically adjusts the shape and path of the convolution kernel to adapt to the target geometry, combining local feature enhancement with topological constraint mechanisms to address the shortcomings of traditional convolution in continuity and adaptability. A kernel deformation strategy similar to "snake walking" is adopted to optimize feature extraction point by point along the target direction, and mathematical tools such as persistent homology are used to constrain the topological continuity of the segmentation results. This significantly improves the ability to extract features of tubular or slender targets, especially showing stronger anti-interference ability in low-contrast and complex backgrounds. After adding dynamic snake convolution to the C3K2 module, it can adaptively process slender structures on the surface of equipment to improve the detection ability of defects, and the effect is more obvious when processing small targets. Figure 3 The C3K2-DSC module in the figure is the C3K2 module with dynamic snake convolution added.

[0166] like Figure 4 As shown in the figure, when C3k is False, the input features pass through the Conv convolution module, the Split module, two Bottleneck-DSC dynamic snake convolution modules, the Concat module, and the Conv convolution module in sequence. The Concat module receives the output features from the first Conv convolution module and the second Bottleneck-DSC dynamic snake convolution module and concatenates them. The Bottleneck-DSC dynamic snake convolution module includes the Input layer, the Conv convolution module, the DysnakeConv dynamic snake convolution module, and the Summation module. The Summation module sums the Input layer features with the DysnakeConv dynamic snake convolution module and outputs the sum.

[0167] When C3k is True, the input features pass through the Conv convolution module, the Split separation module, two C3k-DSC modules, the Concat splicing module, and the Conv convolution module in sequence. Among them, the Concat splicing module receives the output features from the first Conv convolution module and the output features of the second C3k-DSC module respectively for splicing. The C3k-DSC module includes a Conv convolution module, two Bottleneck-DSC bottleneck dynamic snake convolution modules, a Concat splicing module, and a Conv convolution module. The first Conv convolution module also outputs to the Concat splicing module through a Conv convolution module. The Concat splicing module receives the output features and the output features of the second Bottleneck-DSC bottleneck dynamic snake convolution module respectively for splicing.

[0168] Furthermore, a standard two-dimensional convolution with a coordinate of K and a center coordinate of K can be given. i =(x i ,y i ). The coordinates of the 3×3 convolution kernel with a dilation rate of 1 can be expressed as:

[0169] K={(x-1,y-1),(x-1,y),…,(x+1,y+1)}

[0170] Where K is the standard two-dimensional convolution coordinate, x is the horizontal coordinate of the pixel in the horizontal direction of the image, and y is the vertical coordinate of the pixel in the vertical direction of the image, indicating that the center of the convolution kernel is located at the coordinate (x, y), and the eight nearest pixel coordinates are taken into account. In order to give the convolution kernel more flexibility, dynamic snake convolution introduces a learnable deformation offset △ to solve the key problem of standard convolution and variable convolution in segmenting slender tubular structures: when the convolution kernel is freely deformed, its focus area is easy to deviate from the continuous contour of the target structure, resulting in segmentation breakage or offset. In this embodiment, the following can be used Figure 5 The iterative strategy shown selects subsequent positions to be observed for each object to be processed, thereby maintaining the continuity of the region of interest and significantly improving the feature capture capability and segmentation accuracy of slender and complex structures.

[0171] In dynamic snake convolution, the standard convolution kernel is straightened in the X-axis and Y-axis directions. Taking the convolution kernel of size 3×3 and straightened in the X-axis direction as an example, the position of each grid in K can be expressed as K i±c =(x i+c ,y i+c ), c={0, 1, 2, 3, 4} represents the horizontal distance from the center grid. The selection of each grid position in the convolution kernel K is a cumulative process, starting from the center position P i Initially, each grid K i±c The position of depends on the position of the previous grid, K i+1 With K i Compared to the offset Δ = {δ | δ ∈ [-1, 1]}. Therefore, the offset needs to be accumulated to ensure the linear morphology of the convolution kernel. In the figure, the change in the X-axis direction is:

[0172]

[0173] The change in the Y-axis direction is:

[0174]

[0175] like Figure 6As shown, in a specific embodiment, under the changes of the X-axis and Y-axis, the 3×3 convolution covers a 9×9 range after deformation. The dynamic serpentine convolution designed according to the dynamic structure can better adapt to the slender curved structure of the equipment surface, thereby better perceiving the main features.

[0176] The neck network receives feature maps of different scales extracted from the backbone network and further processes and fuses these features to filter, integrate, and enhance the feature information output by the backbone network, thereby improving feature expressiveness and the efficiency of multi-scale information utilization. The neck network can perform operations such as upsampling, downsampling, and feature concatenation on feature maps at different levels, allowing features of different scales to complement and fuse each other, thereby better capturing the characteristic information of objects of different sizes. This preserves high-resolution detail in shallow feature maps for detecting small objects while integrating high-level semantic information in deep feature maps for accurate classification and localization of large objects. Traditional neck networks use an SPPF structure to enhance the receptive field. However, in the case of equipment surface defect detection, small object features are lost and poorly fused. This embodiment improves the neck network by using a feature fusion diffusion pyramid network (FFDPN) with a feature focusing module and feature diffusion mechanism. This improves the output features of each scale with detailed contextual information, enhancing the model's multi-scale feature fusion capabilities and effectively improving the detection of equipment surface defects.

[0177] Feature focusing modules such as Figure 7 As shown, the feature focusing module (FF) receives inputs at three scales: P3 / 8, P4 / 16, and P5 / 32. The P3 layer first undergoes an ADown downsampling module, significantly reducing the number of parameters while maintaining or improving object detection accuracy. The P4 layer performs a 1×1 Conv convolution to maintain a consistent number of channels. The P5 layer first performs an Upsample upsampling operation, followed by a 1×1 Conv convolution. After processing, the features at the three scales are concatenated using a Concat operation to focus the information at each scale. A set of parallel depthwise convolutions is then used to capture rich cross-scale information. Parallel depth convolution includes DWConv5×5, DWConv7×7, DWConv9×9, DWConv11×11 and Identity mapping. After the information is captured, Add is performed to further enhance the input layer feature information, fuse and focus the features of different layers, and use a 1×1 Conv convolution to fuse the channels of features of different scales to obtain features of different scales. After this module, the features of each scale can have detailed contextual information without affecting the local features.

[0178] After feature focusing, Figure 8 The bold arrows show the dynamic diffusion. Figure 8 In the

[15] , the feature focusing module is sequentially connected to the Upsample upsampling layer, the Concat splicing layer, and the C3K2 multi-scale feature extraction layer, and is also sequentially connected to the Conv convolutional layer, the Concat splicing layer, and the C3K2 multi-scale feature extraction layer, thereby obtaining fused features at multiple scales respectively. The fused features at different scales are also spliced ​​with the initially generated focused features to obtain fused features that can be output to the detection head network. Through adaptive weight fusion, the focused high-density semantic information is diffused to each detection layer, ensuring that shallow features retain details while incorporating high-level semantics, thereby solving the problem of missed detection caused by the dilution of information layer by layer in traditional methods. The joint action of the feature focusing module and the feature diffusion mechanism ensures that the features of each scale have detailed contextual information, which is more conducive to subsequent detection operations.

[0179] The traditional detection head network adopts a parallel branch design. Its classification branch only focuses on the probability of the target category, and the localization branch only learns bounding box regression. The two branches process features and generate prediction results through independent convolution. Although this architecture ensures computational efficiency and multi-task independence, the separate design also leads to a lack of information exchange between the two parts. This embodiment designs a dynamic task alignment detection head, improves the detection head network, and realizes the coordinated optimization of classification and positioning through a dual-task interaction mechanism and a dynamic parameter adjustment strategy. It combines task-oriented sample allocation with lightweight design to effectively improve detection accuracy and robustness. Its structure is as follows: Figure 9 shown.

[0180] The three scale features first enter two 3×3 convolution group normalization modules, such as Figure 10 As shown in the figure, the convolution group normalization module includes the use of a group normalization layer for batch normalization and then splicing. It can be understood that the traditional BatchNorm normalizes the batch data of each feature channel by calculating the mean and variance, which depends on the batch size. When the batch is small, the statistical value is unstable; while group normalization divides the feature channels into several groups and normalizes the features within the group without relying on the batch dimension.

[0181] The feature data after feature fusion processing enters the TaskDecompositon task decomposition module. Figure 11 As shown, the task decomposition module adopts a hierarchical structure consisting of N continuous transformation layers, each of which is configured with a nonlinear activation function. Through hierarchical interactive operations between multi-scale features, deep fusion of cross-dimensional feature information is achieved. The task decomposition module dynamically adjusts the contribution of each scale feature through adaptive learning weight parameters w, and finally generates a joint feature representation with strong discriminability, thereby effectively improving the collaborative optimization capability between classification and positioning tasks. In a specific embodiment, the generation formula of the adaptive learning weight is:

[0182]

[0183] Among them, W1 and W2 are learnable weight matrices, σ is the Sigmoid function, the introduction of mean and variance enables weight learning to perceive the distribution characteristics of input features, and the ReLU-Sigmoid activation function enables the weight to focus on differentiated significant features. For shallow high-resolution features, if their distribution variance is low, the weight w i Automatically attenuate and suppress interference; conversely, if a high mean response appears in the deep semantic features, the weight will be increased. Through adaptive weight adjustment, a highly discriminative joint feature representation is ultimately generated, effectively improving the collaborative optimization capabilities between classification and localization tasks.

[0184] The localization branch of the dynamic task alignment detection head (DT-Head) employs variable convolution. Joint features are fed into the Generator Mask & Offset module to generate spatial weight masks and position offset parameters, which are then fed into the DCNv2 module. The DCNv2 module integrates a dynamic modulation mechanism that calculates the offset and weight coefficient for each sampling point in real time based on the spatial distribution of features. By adaptively adjusting the sampling position and response strength of the convolution kernel, it significantly enhances the spatial accuracy of defect localization, effectively overcoming the limitations of traditional fixed convolution kernels in complex defect scenarios. The classification branch of the DT-Head uses joint features for dynamic feature selection for defect classification. To optimize model computational efficiency, the DT-Head employs a shared convolutional architecture, where each detection layer shares the same convolution parameters. This weight sharing strategy significantly reduces model parameter count, effectively reducing memory consumption and achieving lightweight performance while maintaining detection performance. To address the problem of multi-scale object detection, a learnable scale layer is introduced to dynamically adjust the scale of feature maps. This layer optimizes the feature response of objects of different sizes through an adaptive learning mechanism, improving the accuracy and robustness of multi-scale object detection while maintaining model lightweightness.

[0185] This embodiment provides a surface defect detection method that utilizes dynamic snake convolution to improve the C3K2 module. The adaptive convolution kernel deformation strategy can significantly enhance the model's ability to capture slender curved features. A feature fusion diffusion pyramid network, including a feature focusing module and a feature diffusion mechanism, ensures that features at each scale have detailed contextual information, thereby improving the model's multi-scale feature fusion capability. Dynamic task alignment detection heads in the detection head network are used to achieve collaborative optimization of classification and positioning through a dual-task interaction mechanism and a dynamic parameter adjustment strategy. Combined with task-oriented sample allocation and lightweight design, this method improves detection accuracy and robustness.

[0186] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0187] Based on the same inventive concept, the present application also provides a surface defect detection device for implementing the surface defect detection method described above. The solution provided by this device is similar to the solution described in the method described above. Therefore, the specific limitations of one or more surface defect detection device embodiments provided below can be found in the above-mentioned limitations of the surface defect detection method and will not be repeated here.

[0188] In one embodiment, Figure 12 As shown, a surface defect detection device is also provided, and the surface defect detection device includes:

[0189] The pre-processing module 100 is used to extract and fuse features of the surface image to be detected to obtain fused features;

[0190] A fusion module 200 is used to fuse the fusion features based on the adaptive weight parameters to obtain a joint feature;

[0191] The prediction module 300 is used to predict the joint features based on dynamic convolution and shared convolution to obtain positioning prediction results and classification prediction results;

[0192] The adjustment module 400 is used to adjust the positioning prediction results and the classification prediction results based on the learnable scale parameters to obtain the defect detection results.

[0193] In one embodiment, the pre-processing module 100 is further configured to:

[0194] Performing feature extraction on the surface image to be detected based on a pre-trained backbone network to obtain image extraction features of multiple different scales;

[0195] Based on the pre-trained neck network, feature fusion is performed on the plurality of image extraction features to obtain the fused feature.

[0196] In one embodiment, the backbone network includes multiple multi-scale feature extraction modules, which are used to: perform dynamic convolution processing based on input features to obtain multi-scale convolution features; and perform splicing based on the multi-scale convolution features and the input features to obtain output features;

[0197] The multi-scale feature extraction module includes a variable convolution module, which is used to: determine the target geometric direction of the input feature based on the input feature and the persistent homology tool; determine the offset parameter based on multiple sampling points of the input feature and the target geometric direction; determine the matching convolution kernel and sampling path based on the offset parameter; and perform dynamic convolution processing on the input feature based on the convolution kernel and the sampling path to obtain the multi-scale convolution feature.

[0198] In one embodiment, determining the offset parameter based on the plurality of sampling points of the input feature and the target geometric orientation includes:

[0199] Based on each of the sampling points, determining an offset corresponding to each sampling point;

[0200] Based on the target geometric direction of the input feature, the offsets of the plurality of sampling points are accumulated to obtain the offset parameter.

[0201] In one embodiment, the neck network includes a feature fusion diffusion pyramid network for:

[0202] Performing multi-scale processing and parallel depth convolution based on the features extracted from the multiple images to obtain high-density semantic information;

[0203] Based on an adaptive weight parameter, the high-density semantic information is weightedly superimposed with the plurality of image extraction features to obtain a plurality of fusion features.

[0204] In one embodiment, the feature fusion diffusion pyramid network includes a feature focusing module, the feature focusing module includes a multi-scale processing module and a parallel depth convolution module, and the multi-scale processing module is used to:

[0205] Based on the scale of the features extracted from each image, feature processing is performed to match the scale to obtain a preset scale feature; the feature processing includes at least one of upsampling, downsampling, and convolution;

[0206] splicing a plurality of the preset scale features to obtain a focus feature;

[0207] The focused features are input into the parallel depthwise convolution module to obtain the high-density semantic information.

[0208] In one embodiment, the prediction module 300 is further configured to:

[0209] Generating geometric adaptability parameters based on the joint features; adjusting the sampling position and weight of the convolution kernel based on the geometric adaptability parameters to obtain a positioning prediction result;

[0210] The classification weight is dynamically adjusted based on the response feature of the joint feature to obtain a classification prediction result.

[0211] In one embodiment, the geometric adaptability parameters include a spatial weight mask and a position offset parameter, and the prediction module 300 is further configured to:

[0212] performing weighted processing on the sampling points in the joint feature based on the spatial weight mask;

[0213] Adjusting the sampling position of the variable convolution kernel based on the position offset parameter;

[0214] The joint features are processed based on the adjusted convolution kernel to obtain a positioning prediction result.

[0215] In one embodiment, the prediction module 300 is further configured to:

[0216] Dimensionally compressing the joint features based on a shared convolution structure to obtain compressed features;

[0217] Based on the response strength of the compressed features, a learnable response scaling layer is used to dynamically adjust the classification weights of multi-scale objects;

[0218] Based on the adjusted classification weights, a classification prediction result is determined; the classification prediction result includes a defect category and a confidence level corresponding to the defect category.

[0219] Each module in the surface defect detection device described above can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0220] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 13As shown. The computer device includes a processor, memory, a communication interface, a display screen, and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal via wired or wireless communication. The wireless communication can be achieved via Wi-Fi, a mobile cellular network, NFC (near-field communication), or other technologies. When executed by the processor, the computer program implements a surface defect detection method. The display screen of the computer device can be a liquid crystal display or an electronic ink display. The input device of the computer device can be a touch layer covering the display screen, or keys, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse.

[0221] Those skilled in the art will understand that Figure 13 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0222] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the surface defect detection method of any of the above embodiments is implemented:

[0223] Acquiring an image of a surface to be inspected captured by an image sensor;

[0224] The surface image to be inspected is input into a pre-trained defect detection model to obtain a defect detection result; the defect detection model includes a backbone network, a neck network and a detection head network; the neck network includes a feature fusion diffusion pyramid network, which is used to: perform multi-scale processing and parallel deep convolution based on multiple image extraction features of different scales output by the backbone network to obtain high-density semantic information; based on adaptive weights, the high-density semantic information is weightedly superimposed with multiple image extraction features to obtain multiple enhanced fusion features; the detection head network is used to determine the defect detection result based on the multiple enhanced fusion features.

[0225] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the surface defect detection method of any of the above embodiments is implemented:

[0226] Acquiring an image of a surface to be inspected captured by an image sensor;

[0227] The surface image to be inspected is input into a pre-trained defect detection model to obtain a defect detection result; the defect detection model includes a backbone network, a neck network and a detection head network; the neck network includes a feature fusion diffusion pyramid network, which is used to: perform multi-scale processing and parallel deep convolution based on multiple image extraction features of different scales output by the backbone network to obtain high-density semantic information; based on adaptive weights, the high-density semantic information is weightedly superimposed with multiple image extraction features to obtain multiple enhanced fusion features; the detection head network is used to determine the defect detection result based on the multiple enhanced fusion features.

[0228] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0229] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.

[0230] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0231] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A surface defect detection method, characterized in that: The surface defect detection method comprises: Perform feature extraction and feature fusion on the surface image to be detected to obtain fusion features; fusing the fused features based on an adaptive weight parameter to obtain a joint feature; Based on dynamic convolution and shared convolution, the joint features are predicted respectively to obtain positioning prediction results and classification prediction results; Based on the learnable scale parameter, the positioning prediction result and the classification prediction result are adjusted to obtain a defect detection result.

2. The surface defect detection method according to claim 1, characterized in that: The feature extraction and feature fusion of the surface image to be detected to obtain the fusion feature includes: Performing feature extraction on the surface image to be detected based on a pre-trained backbone network to obtain image extraction features of multiple different scales; Based on the pre-trained neck network, feature fusion is performed on the plurality of image extraction features to obtain the fused feature.

3. The surface defect detection method according to claim 2, characterized in that: The backbone network includes multiple multi-scale feature extraction modules, which are used to: perform dynamic convolution processing based on input features to obtain multi-scale convolution features; and perform splicing based on the multi-scale convolution features and the input features to obtain output features; The multi-scale feature extraction module includes a variable convolution module, which is used to: determine the target geometric direction of the input feature based on the input feature and the persistent homology tool; determine an offset parameter based on multiple sampling points of the input feature and the target geometric direction; and determine a matching convolution kernel and sampling path based on the offset parameter; Based on the convolution kernel and the sampling path, dynamic convolution processing is performed on the input feature to obtain the multi-scale convolution feature.

4. The surface defect detection method according to claim 3, characterized in that: Determining the offset parameters based on the multiple sampling points of the input features and the target geometric direction includes: Based on each of the sampling points, determining an offset corresponding to each sampling point; Based on the target geometric direction of the input feature, the offsets of the plurality of sampling points are accumulated to obtain the offset parameter.

5. The surface defect detection method according to claim 2, characterized in that: The neck network includes a feature fusion diffusion pyramid network for: Performing multi-scale processing and parallel depth convolution based on the features extracted from the multiple images to obtain high-density semantic information; Based on an adaptive weight parameter, the high-density semantic information is weightedly superimposed with the plurality of image extraction features to obtain a plurality of fusion features.

6. The surface defect detection method according to claim 5, characterized in that: The feature fusion diffusion pyramid network includes a feature focusing module, which includes a multi-scale processing module and a parallel depth convolution module. The multi-scale processing module is used to: Based on the scale of the features extracted from each image, feature processing is performed to match the scale to obtain a preset scale feature; the feature processing includes at least one of upsampling, downsampling, and convolution; splicing a plurality of the preset scale features to obtain a focus feature; The focused features are input into the parallel depthwise convolution module to obtain the high-density semantic information.

7. The surface defect detection method according to claim 1, characterized in that: The joint features are predicted based on dynamic convolution and shared convolution respectively to obtain positioning prediction results and classification prediction results, including: Generating geometric adaptability parameters based on the joint features; adjusting the sampling position and weight of the convolution kernel based on the geometric adaptability parameters to obtain a positioning prediction result; The classification weight is dynamically adjusted based on the response feature of the joint feature to obtain a classification prediction result.

8. The surface defect detection method according to claim 7, characterized in that: The geometric adaptability parameters include a spatial weight mask and a position offset parameter. Adjusting the sampling position and weight of the convolution kernel based on the geometric adaptability parameters to obtain a positioning prediction result includes: performing weighted processing on the sampling points in the joint feature based on the spatial weight mask; Adjusting the sampling position of the variable convolution kernel based on the position offset parameter; The joint features are processed based on the adjusted convolution kernel to obtain a positioning prediction result.

9. The surface defect detection method according to claim 7, characterized in that: The dynamically adjusting the classification weight based on the response feature of the joint feature to obtain the classification prediction result includes: Dimensionally compressing the joint features based on a shared convolution structure to obtain compressed features; Based on the response strength of the compressed features, a learnable response scaling layer is used to dynamically adjust the classification weights of multi-scale objects; Based on the adjusted classification weights, a classification prediction result is determined; the classification prediction result includes a defect category and a confidence level corresponding to the defect category.

10. A surface defect detection device, characterized in that: The surface defect detection device comprises: A preprocessing module is used to extract and fuse features of the surface image to be detected to obtain fused features; A fusion module, configured to fuse the fusion features based on an adaptive weight parameter to obtain a joint feature; A prediction module, configured to predict the joint features based on dynamic convolution and shared convolution, respectively, to obtain a positioning prediction result and a classification prediction result; An adjustment module is used to adjust the positioning prediction result and the classification prediction result based on a learnable scale parameter to obtain a defect detection result.

Citation Information

Patent Citations

  • Method for realizing BPCB surface defect detection based on CNN

    CN115409784A

  • Turbine starter appearance defect detection method based on multi-scale feature extraction

    CN119625404A

  • Welded part surface defect detection method and system based on gray level co-occurrence matrix and YOLOv11

    CN120495243A

Cited By

  • Seamless steel tube surface defect positioning and detecting method and system based on deep learning

    CN121120628A

  • Drainage pipeline defect detection method and device based on cost-sensitive risk decision and scale perception optimization, and computer readable storage medium

    CN122049873A