An industrial defect detection method based on convolutional neural network
Patent Information
- Application Number
- CN202610765893.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-29
- Publication Date
- 2026-09-25
AI Technical Summary
上述各向同性的均匀偏移策略导致沿运动轴的采样点逆向回拉幅度不足、去模糊残留拖尾,同时垂直于运动轴的采样点被施加了不必要的大幅偏移而引入横向伪影,两个正交维度上的定位精度均受到损害
[0011]与现有技术相比,本申请提出一种基于卷积神经网络的工业缺陷检测方法。其通过在骨干网络浅层提取保留模糊纹理的缺陷语义特征后,利用并行的运动矢量先验感知分支解算出逐像素的运动方向角度与模糊尺度参数,进而基于各采样点网格坐标与运动传播轴之间的方向投影耦合关系,构建逐采样点差异化的各向异性反向空间偏移矩阵,驱动可变形卷积的采样网格呈现沿运动轴强收敛、垂直运动轴弱扰动的椭圆形分布,在特征空间内完成弥散能量的精准逆向聚合与边缘锐化。锐化后的特征经由双向特征金字塔进行多尺度语义融合,最终通过融合边界不确定性方差的解耦检测头动态惩罚高抖动预测框,同时提升运动方向上的去模糊彻底性与正交方向上的边缘保真度,实现高速场景下缺陷的高精度分类与稳定边界定位。
Smart Images

Figure CN122820550A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of computer vision and deep learning technology, and more specifically, to an industrial defect detection method based on convolutional neural networks. Background Technology
[0002] In modern high-speed industrial production lines, workpieces pass through the field of view of a linear array camera at extremely high linear speeds, inevitably resulting in linear trailing blur along the direction of motion in the captured images. Although convolutional neural network-based defect detection methods have become the mainstream technology in industrial quality inspection, their bounding box regression mechanism is highly dependent on the high-frequency spatial gradient information of the defect edges. The feature diffusion caused by motion blur leads to physical drift of the true defect boundary, directly causing a serious degradation in detection and positioning accuracy.
[0003] In existing technologies, some schemes attempt to introduce deformable convolution mechanisms constrained by motion vectors during the feature extraction stage. This involves constructing an inverse spatial offset matrix by predicting the local motion direction and blur scale parameters, guiding the convolution sampling points to converge and resample in the opposite direction of the blur to eliminate diffusion at the feature level. However, such schemes apply a uniform inverse offset to all sampling points in the convolution kernel when generating the offset matrix, implicitly assuming that each sampling point contributes isotropically to the motion blur. In reality, motion blur energy diffusion exhibits strict directional anisotropy; pixel energy only undergoes linear tailing diffusion along the motion direction, while high-frequency edge information perpendicular to the motion direction remains almost undamaged. This isotropic uniform offset strategy results in insufficient inverse pullback amplitude for sampling points along the motion axis, leaving residual tailing after blurring. Simultaneously, sampling points perpendicular to the motion axis are subjected to unnecessary large offsets, introducing lateral artifacts, thus compromising positioning accuracy in both orthogonal dimensions.
[0004] Therefore, an optimized industrial defect detection method based on convolutional neural networks is desired. Summary of the Invention
[0005] To address the aforementioned technical problems, this application is proposed. Embodiments of this application provide an industrial defect detection method based on a convolutional neural network, comprising:
[0006] Step 1: Extract shallow fuzzy defect features from the motion-blurred grayscale image of the production line acquired by an industrial line scan camera on a high-speed production line to obtain the shallow defect semantic feature tensor.
[0007] Step 2: Perform defect kinematic feature solving based on local perception branch on the shallow defect semantic feature tensor to obtain the local motion vector field map;
[0008] Step 3: Based on the motion direction angle and fuzzy scale parameters encoded in the local motion vector field map, perform feature-level inverse aggregation defuzzification on the shallow defect semantic feature tensor based on motion dynamics constraints to obtain the inverse aggregation sharpened feature tensor.
[0009] Step 4: Through the top-down semantic transmission path and bottom-up spatial localization path of the feature pyramid network, cross-scale feature alignment and multi-scale semantic fusion are performed on the inverse aggregated sharpening feature tensor to obtain a multi-scale defect feature pyramid.
[0010] Step 5: Based on the decoupled detection head, perform defect category and location prediction on the multi-scale defect feature pyramid to output a defect localization detection result set.
[0011] Compared with existing technologies, this application proposes an industrial defect detection method based on convolutional neural networks. After extracting defect semantic features that retain blurred textures in the shallow layers of the backbone network, it utilizes parallel motion vector prior perception branches to calculate the pixel-by-pixel motion direction angle and blur scale parameters. Then, based on the directional projection coupling relationship between the grid coordinates of each sampling point and the motion propagation axis, it constructs a differentially distributed anisotropic inverse spatial offset matrix for each sampling point. This drives the deformable convolutional sampling grid to exhibit an elliptical distribution with strong convergence along the motion axis and weak perturbation perpendicular to the motion axis, achieving precise inverse aggregation of diffuse energy and edge sharpening within the feature space. The sharpened features undergo multi-scale semantic fusion via a bidirectional feature pyramid. Finally, by decoupling the detection head through the fusion boundary uncertainty variance, it dynamically penalizes high-jitter prediction boxes, simultaneously improving the thoroughness of deblurring in the motion direction and the edge fidelity in the orthogonal direction, achieving high-precision defect classification and stable boundary localization in high-speed scenes. Attached Figure Description
[0012] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0013] Figure 1 This is a flowchart of an industrial defect detection method based on a convolutional neural network according to an embodiment of this application;
[0014] Figure 2 This is a schematic diagram of data flow in an industrial defect detection method based on a convolutional neural network according to an embodiment of this application.
[0015] Figure 3This is a flowchart illustrating a method for industrial defect detection based on a convolutional neural network according to an embodiment of this application, which involves solving the kinematic features of shallow defects based on a local perceptual branch to obtain a local motion vector field map.
[0016] Figure 4 This document describes a flowchart illustrating a convolutional neural network-based industrial defect detection method according to an embodiment of this application. It describes a process for performing feature-level reverse aggregation deblurring on shallow defect semantic feature tensors based on motion direction angles and fuzzy scale parameters encoded in a local motion vector field map to obtain a reverse aggregation sharpened feature tensor.
[0017] Figure 5 This is a flowchart illustrating the determination of the inverse spatial offset matrix based on the motion direction tensor and motion scale tensor in an industrial defect detection method based on a convolutional neural network according to an embodiment of this application.
[0018] Figure 6 This is a flowchart illustrating a method for industrial defect detection based on convolutional neural networks, according to an embodiment of this application, which uses a decoupled detection head to predict the defect category and location of a multi-scale defect feature pyramid to output a defect location detection result set. Detailed Implementation
[0019] Hereinafter, exemplary embodiments according to this application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments of this application. It should be understood that this application is not limited to the exemplary embodiments described herein.
[0020] As indicated in this application and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not specifically singular and may include plural forms. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.
[0021] While this application makes various references to certain modules of the systems according to embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The modules described are merely illustrative, and different aspects of the systems and methods may use different modules.
[0022] Flowcharts are used in this application to illustrate the operations performed by the system according to embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, various steps can be processed in reverse order or simultaneously as needed. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.
[0023] In high-speed industrial production line defect detection scenarios, existing convolutional neural network-based detection methods suffer from significant degradation in bounding box regression accuracy when dealing with motion-blurred images due to severe dispersion of high-frequency gradient information at defect edges along the motion direction. While some solutions introduce deformable convolution for feature-level inverse convergence, applying a uniform offset to all sampling points ignores the anisotropic nature of motion blur energy dispersion, resulting in incomplete deblurring along the motion axis and the introduction of lateral artifacts in the vertical direction. Therefore, this application proposes an industrial defect detection method based on convolutional neural networks. This method first extracts shallow features from the motion-blurred grayscale image of the pipeline through a backbone network, preserving blurred texture and defect semantic information. Then, it uses a parallel motion vector perception branch to calculate the motion direction angle and blur scale parameters pixel by pixel, generating a local motion vector field map. Based on this, it calculates the directional projection coupling weight between the grid coordinates of each sampling point and the motion propagation axis, and uses this weight to perform differential amplitude modulation on the basic inverse offset vector for each sampling point. This results in strong inverse convergence strength for sampling points along the motion axis while only minimal perturbation in the vertical direction, generating an anisotropic inverse spatial offset matrix to drive deformable convolution to complete feature-level accurate deblurring. The sharpened features are then subjected to multi-scale semantic fusion through a bidirectional feature pyramid. Finally, the decoupled detection head dynamically penalizes high-jitter prediction boxes by fusing boundary uncertainty variance and outputs the defect localization detection results. This method ensures thorough deblurring in the motion direction and edge fidelity in the orthogonal direction without increasing additional inference overhead.
[0024] Figure 1 This is a flowchart of an industrial defect detection method based on a convolutional neural network according to an embodiment of this application. Figure 2 This is a schematic diagram of the data flow in an industrial defect detection method based on a convolutional neural network according to an embodiment of this application. Figure 1 and Figure 2As shown, an industrial defect detection method based on a convolutional neural network according to an embodiment of this application includes: S1, extracting shallow-layer blurred defect features from a motion-blurred grayscale image of an industrial linear array camera acquired on a high-speed production line to obtain a shallow-layer defect semantic feature tensor; S2, performing defect kinematic feature calculation based on a local perception branch on the shallow-layer defect semantic feature tensor to obtain a local motion vector field map; S3, performing feature-level inverse aggregation deblurring based on motion dynamics constraints on the shallow-layer defect semantic feature tensor according to the motion direction angle and blur scale parameters encoded in the local motion vector field map to obtain an inverse aggregation sharpened feature tensor; S4, performing cross-scale feature alignment and multi-scale semantic fusion on the inverse aggregation sharpened feature tensor through the top-down semantic transfer path and bottom-up spatial positioning path of the feature pyramid network to obtain a multi-scale defect feature pyramid; S5, predicting the defect category and location of the multi-scale defect feature pyramid based on a decoupled detection head to output a defect localization detection result set.
[0025] Specifically, in step S1, shallow blur defect features are extracted from the motion-blurred grayscale image of the production line acquired by an industrial line scan camera on a high-speed production line to obtain a shallow defect semantic feature tensor. It should be noted that, since the edge gradient information of the defect region in the grayscale image acquired by the industrial line scan camera on the high-speed production line has diffused along the motion direction, subsequent motion vector calculations and reverse aggregation operations rely on intermediate feature representations that retain the blurred texture distribution as the computational basis. Based on this, the technical solution of this application first extracts shallow blur defect features from the motion-blurred grayscale image of the production line to obtain a shallow defect semantic feature tensor. Through the above processing, the grayscale gradient distribution and the diffuse pixel distribution along the motion direction of the defect region can be preserved while compressing spatial redundancy, providing feature inputs with both spatial topological structure and blurred prior information for subsequent branches.
[0026] More specifically, in a concrete example of this application, a sliding convolution operation is first performed on the pipeline motion-blurred grayscale image using a 7×7 large receptive field 2D convolution kernel with a stride of 2. The 7×7 kernel size can cover a large spatial neighborhood in a single convolution operation, thereby capturing cross-pixel grayscale gradient changes caused by motion blur. Setting the stride to 2 reduces the spatial resolution of the output feature map to half that of the input, completing spatial compression in the initial stage to reduce subsequent computational load. After the convolution operation is completed, batch normalization is applied to the output to stabilize the numerical distribution range of each channel. Subsequently, a nonlinear mapping is performed using the ReLU activation function to truncate negative responses and retain positive high-frequency gradient activation signals, resulting in an initial high-frequency response feature map.
[0027] Secondly, the initial high-frequency response feature map is input into the residual downsampling module. This module contains two parallel paths: the main branch uses a 3×3 convolutional kernel for local topological feature extraction and a spatial downsampling operation with a stride of 2 to further expand the receptive field and aggregate the structural relationships between adjacent pixels; the identity mapping branch uses a 1×1 convolutional kernel to align the input with the number of channels and spatial resolution. The outputs of the two branches are element-wise summed and then nonlinearly activated to obtain the intermediate-level topological aggregated feature tensor. The introduction of residual connections ensures the effective flow of gradients during backpropagation, avoiding feature degradation caused by the increase in the number of network layers.
[0028] Finally, a 1×1 pointwise convolution operator is used to perform a cross-channel linear weighted combination of the intermediate topological aggregation feature tensor. Since different convolutional kernels in the preceding steps respond to different morphological features of the defects and different diffusion modes of motion blur, the 1×1 convolution adjusts the weight ratio of each channel, thus re-integrating these discrete feature responses into a unified semantic expression. After activation, this outputs a shallow defect semantic feature tensor. This tensor preserves the relative positional topological relationships of the defect region in the spatial dimension and encodes the combined semantic information of blurred texture and edge gradients in the channel dimension.
[0029] Specifically, in step S2, the shallow defect semantic feature tensor is processed using defect kinematic features based on a local perceptual branch to obtain a local motion vector field map. It should be noted that, given that the subsequent feature-level inverse aggregation deblurring operation requires explicit knowledge of the specific physical direction and diffusion amplitude of defect feature dispersion at each spatial location in order to construct a reverse spatial offset matrix with physical geometric constraints to guide the sampling point distribution of deformable convolution, and although the shallow defect semantic feature tensor implicitly contains prior information on the directional texture of motion blur, this information has not yet been explicitly computed into motion parameters that can be directly consumed by the offset matrix. Based on this, the technical solution of this application further processes the shallow defect semantic feature tensor using defect kinematic features based on a local perceptual branch to obtain a local motion vector field map. Through the above processing, the implicit directional texture of the feature tensor can be transformed into an explicit parameterized expression that encodes the motion direction angle and blur scale pixel by pixel, providing a quantitative physical constraint basis for the subsequent generation of the anisotropic offset matrix.
[0030] Figure 3 This is a flowchart illustrating a method for industrial defect detection based on a convolutional neural network, according to an embodiment of this application, which involves calculating the kinematic features of shallow defect semantic feature tensors using a local perception branch to obtain a local motion vector field map. Figure 3As shown, step S2 includes: S21, performing motion blur prior enhancement based on frequency domain transformation on the shallow defect semantic feature tensor to obtain a frequency domain enhanced feature tensor; S22, performing pixel-wise dual-track regression of motion direction and blur scale on the frequency domain enhanced feature tensor through two sets of parallel 1×1 convolution kernels to obtain a motion parameter mapping tensor; S23, performing spatial consistency smoothing of motion parameters between adjacent pixels on the motion parameter mapping tensor through a 3×3 depth-separable convolution operator to obtain a local motion vector field map.
[0031] In step S21, motion fuzz prior enhancement based on frequency domain transformation is performed on the shallow defect semantic feature tensor to obtain a frequency domain enhanced feature tensor. It should be noted that since the motion fuzz directionality information in the shallow defect semantic feature tensor exists in the spatial domain as an implicit texture distribution, direct regression calculation is easily affected by background noise, reducing the accuracy of motion parameter estimation. Therefore, the technical solution of this application further performs motion fuzz prior enhancement based on frequency domain transformation on the shallow defect semantic feature tensor to obtain a frequency domain enhanced feature tensor. Through the above processing, the frequency component signal related to the motion fuzz direction can be directionally amplified in the frequency domain space, suppressing irrelevant background noise and providing a feature input with a higher signal-to-noise ratio for subsequent motion direction and fuzz scale regression.
[0032] More specifically, in a concrete example of this application, the frequency domain transformation enhancement process is as follows. First, a two-dimensional fast Fourier transform is performed along the spatial dimension of the shallow defect semantic feature tensor, mapping the spatial domain features of each channel to a frequency domain complex feature spectrum. In the frequency domain representation, the directional energy dispersion generated by motion blur manifests as a concentration of frequency amplitudes distributed along a specific angle. This prior information is difficult to directly observe in the spatial domain, but it is presented as an explicit, structured distribution in the frequency domain. Subsequently, the frequency domain complex feature spectrum is multiplied element-wise with a learnable frequency domain attention mask. This mask adaptively learns the amplification weights for the motion blur direction-related frequency components and the suppression weights for background spurious frequency components during network training. Finally, a two-dimensional inverse fast Fourier transform is performed on the modulated frequency domain feature spectrum to transform it back to the spatial domain, and the element-wise residuals are added to the original shallow defect semantic feature tensor to obtain the frequency domain enhanced feature tensor. The introduction of residual connections ensures that the original spatial semantic information is not destroyed, while simultaneously superimposing the motion prior enhancement signal after selective amplification in the frequency domain.
[0033] In step S22, the motion parameter mapping tensor is obtained by performing pixel-by-pixel dual-track regression of motion direction and fuzziness scale on the frequency domain enhanced feature tensor using two sets of parallel 1×1 convolution kernels. It should be noted that since the frequency domain enhanced feature tensor already contains selectively amplified motion fuzziness directionality prior signals, but the motion direction angle and fuzziness scale are two types of parameters with different physical meanings—the former representing the spatial orientation of the fuzziness and the latter representing the fuzziness's spread—they need to be solved separately through independent regression paths to avoid gradient interference. Based on this, the technical solution of this application further uses two sets of parallel 1×1 convolution kernels to perform pixel-by-pixel dual-track regression of motion direction and fuzziness scale on the frequency domain enhanced feature tensor to obtain the motion parameter mapping tensor. Through the above processing, the high-dimensional features after frequency domain enhancement can be transformed into pixel-by-pixel motion parameter codes with clear physical semantics, providing directional and magnitude constraints for the subsequent construction of the offset matrix.
[0034] More specifically, in a concrete example of this application, the dual-track regression process is implemented as follows: The frequency domain enhanced feature tensor is simultaneously fed into two parallel regression paths. In the first regression path, a first set of 1×1 convolutional kernels is used to perform cross-channel linear dimensionality reduction on the frequency domain enhanced feature tensor, compressing the multi-channel features into a single-channel angle prediction value. Then, the Tanh activation function maps the output to a range of -1 to +1, and multiplies it by a constant π, so that the final output motion direction angle value falls within the physical radian domain of -π to +π, corresponding to any directional blur angle that may be generated by the movement of the workpiece on the assembly line. In the second regression path, a second set of 1×1 convolutional kernels is used for the same cross-channel linear dimensionality reduction. Then, the Sigmoid activation function maps the output to a range of zero to one, and multiplies it by a preset maximum blur scale constant, so that the final output blur scale value represents the maximum number of swept pixels of motion diffusion at that pixel location. Both paths use 1×1 convolutional kernels to achieve independent calculation per pixel, ensuring the decoupling of parameter estimation between spatial locations. Finally, the motion direction angle mapping map output by the first path and the fuzzy scale mapping map output by the second path are concatenated along the channel dimension to obtain the motion parameter mapping tensor. This tensor encodes two physical parameters, the motion direction and the diffusion amplitude, of the corresponding pixel at each spatial location.
[0035] In step S23, a 3×3 depth-separable convolution operator is used to perform spatial consistency smoothing of the motion parameters between adjacent pixels on the motion parameter mapping tensor to obtain a local motion vector field map. It should be noted that, since the motion parameter mapping tensor obtained from pixel-by-pixel independent regression may contain isolated noise jumps in the motion direction angles and fuzzy scale values between adjacent pixels, while the physical motion of workpieces in a real assembly line scenario exhibits high continuity and consistency within the local spatial neighborhood. Based on this, the technical solution of this application further uses a 3×3 depth-separable convolution operator to perform spatial consistency smoothing of the motion parameters between adjacent pixels on the motion parameter mapping tensor to obtain a local motion vector field map. Through the above processing, outlier vectors and isolated noise generated in pixel-by-pixel regression can be eliminated, making the output motion vector field exhibit a smooth distribution in space that conforms to the constraints of physical motion continuity.
[0036] More specifically, in a concrete example of this application, the implementation process of spatially consistent smoothing is as follows: The motion parameter mapping tensor is fed into a 3×3 depthwise separable convolution operator. The depthwise separable convolution consists of two stages connected in series: depthwise convolution and pointwise convolution. In the depthwise convolution stage, a 3×3 spatial convolution kernel is independently applied to each channel of the motion parameter mapping tensor, causing the motion parameter value at each pixel location to be weighted and averaged with the corresponding parameter values of its eight neighboring pixels. This achieves smoothing constraints on the motion direction and fuzzy scale within the local spatial neighborhood while maintaining parameter independence between channels. In the pointwise convolution stage, a 1×1 convolution kernel is used to perform cross-channel information exchange on the smoothed results of each channel, establishing a cooperative constraint relationship between the motion direction channel and the fuzzy scale channel. After convolution, a LeakyReLU activation function is applied to the output for nonlinear mapping, introducing nonlinearity while preserving the expressive power of negative angular features, resulting in a local motion vector field map. Compared to standard convolution, depthwise separable convolution has fewer parameters and lower computational cost, making it suitable for this branch as a lightweight auxiliary perception path without adding extra burden to the overall network's inference efficiency.
[0037] Specifically, in step S3, based on the motion direction angle and fuzzy scale parameters encoded in the local motion vector field map, the shallow defect semantic feature tensor is deblurred using feature-level inverse aggregation based on motion dynamics constraints to obtain an inverse aggregation sharpened feature tensor. It should be noted that, given that the high-frequency gradient information of the defect edges in the shallow defect semantic feature tensor has undergone linear tailing and diffusion along the motion direction, and the local motion vector field map has explicitly encoded the motion direction angle and fuzzy scale parameters at each spatial location, it provides a quantitative basis for guiding feature inverse convergence from a physical geometry perspective. This motion prior needs to be utilized to directly complete the inverse aggregation of diffused energy in the feature space, avoiding the destruction of high-order texture semantics crucial for defect classification by traditional pixel-level image deblurring schemes that prioritize reconstruction error as the optimization objective. Based on this, the technical solution of this application further performs feature-level inverse aggregation deblurring based on motion dynamics constraints on the shallow defect semantic feature tensor using the motion direction angle and fuzzy scale parameters encoded in the local motion vector field map to obtain an inverse aggregation sharpened feature tensor. Through the above processing, the defect edge energy diffused along the motion direction can be reversed and converged back to its original spatial position in the feature high-dimensional space, reconstructing a feature expression with sharp edges and distinct contrast. At the same time, the high-frequency edge gradient information that is not damaged in the direction of motion is preserved, providing a feature benchmark with high positioning accuracy in both orthogonal dimensions for the subsequent bounding box regression of the detection head.
[0038] Figure 4 This document describes a flowchart illustrating a convolutional neural network-based industrial defect detection method according to an embodiment of this application. It describes a process for deblurring shallow defect semantic feature tensors based on feature-level inverse aggregation constrained by motion dynamics constraints, using motion direction angles and fuzzy scale parameters encoded in a local motion vector field map to obtain an inversely aggregated sharpened feature tensor. (See flowchart for example.) Figure 4 As shown, step S3 includes: S31, performing feature slicing and decoupling on the local motion vector field map along the channel dimension to obtain the motion direction tensor and the motion scale tensor; S32, determining the inverse spatial offset matrix based on the motion direction tensor and the motion scale tensor; S33, performing anisotropic inverse convergence resampling and edge feature weighted sharpening on the shallow defect semantic feature tensor based on the inverse spatial offset matrix to obtain the inverse aggregated sharpening feature tensor.
[0039] In step S31, the local motion vector field map is feature-sliced and decoupled along the channel dimension to obtain the motion direction tensor and the motion scale tensor. It should be noted that since the local motion vector field map encodes two types of parameters with different physical meanings—motion direction angle and fuzzy scale—along the channel dimension, and the subsequent construction of the inverse spatial offset matrix requires applying trigonometric function projection to the direction information and scalar scaling to the scale information, the computational paths of these two are independent. Based on this, the technical solution of this application further performs feature-slicing and decoupling along the channel dimension to obtain the motion direction tensor and the motion scale tensor. Through the above processing, the hybrid encoded motion parameters can be decomposed into two unambiguous tensors that can independently participate in subsequent geometric calculations, providing explicit data inputs for the direction constraint and magnitude constraint of the inverse offset matrix.
[0040] More specifically, in a concrete example of this application, the channel decoupling process is as follows: A slicing operation is performed on the local motion vector field map along its channel dimension at preset index positions. The feature slices corresponding to the first half of the channels are extracted as motion direction tensors. The value at each spatial position in this tensor represents the physical motion angle of the defect feature dispersion at that pixel, with a value range covering the complete radian domain from negative π to positive π. The feature slices corresponding to the second half of the channels are extracted as motion scale tensors. The value at each spatial position in this tensor represents the maximum affected pixel range of motion blur at that pixel. The slicing operation itself does not introduce any learnable parameters; it only performs deterministic tensor splitting based on the channel index, ensuring zero computational overhead and numerical losslessness in the decoupling process. After decoupling, the motion direction tensor will be used in subsequent steps to calculate the inverse unit direction component, and the motion scale tensor will be used to control the magnitude of the inverse offset vector. The two work together to generate the inverse spatial offset matrix.
[0041] In step S32, the reverse spatial offset matrix is determined based on the motion direction tensor and the motion scale tensor. It should be noted that, given that the motion direction tensor and the motion scale tensor have respectively encoded the physical orientation and maximum sweep range of the defect feature diffusion at each spatial location, these two independent kinematic parameters need to be transformed into a unified data structure that can directly drive the deformable convolution operator for inverse convergence sampling. Furthermore, this transformation process must fully consider the directional anisotropy of motion blur energy diffusion, ensuring that sampling points in different spatial orientations receive differentiated offsets that match their degree of blur influence, rather than applying a uniform inverse offset to all sampling points. Based on this, the technical solution of this application further determines the reverse spatial offset matrix based on the motion direction tensor and the motion scale tensor. Through the above processing, an offset field that accurately reflects the anisotropic physical nature of motion blur direction can be generated, so that the sampling points arranged along the motion axis have strong inverse convergence force to eliminate diffusion tail, and the sampling points perpendicular to the motion axis are only subject to minimal perturbation to retain undamaged high-frequency edge information. This provides sampling coordinate guidance that is physically reasonable in both orthogonal dimensions for subsequent inverse convergence resampling of anisotropic deformable convolution.
[0042] Figure 5 This is a flowchart illustrating the determination of the inverse spatial offset matrix based on the motion direction tensor and motion scale tensor in an industrial defect detection method based on a convolutional neural network according to an embodiment of this application. Figure 5 As shown, step S32 includes: S321, constructing a normalized motion direction unit vector based on the angle values in the motion direction tensor, performing inner product projection and normalization on the coordinate vectors of each sampling point of the convolutional grid and the motion direction unit vector to obtain the direction projection coupling weight matrix; S322, using the direction projection coupling weight matrix as a per-sampling-point modulation factor, performing differential amplitude modulation on the basic offset vector after the motion direction tensor is scaled by the inverse triangular projection and the motion scale tensor to generate an anisotropic inverse offset vector set; S323, stacking the differential offset vectors of each sampling point in the anisotropic inverse offset vector set according to the sampling point index order and rearranging them in the deformable convolutional standard offset field format to obtain the inverse spatial offset matrix.
[0043] In step S321, a normalized motion direction unit vector is constructed based on the angle values in the motion direction tensor. This is then used to perform an inner product projection and normalization on the coordinate vectors of each sampling point in the convolutional grid with the motion direction unit vector to obtain the direction projection coupling weight matrix. It should be noted that, given the strict directional anisotropy of motion blur energy dispersion, pixel energy only undergoes linear tail diffusion along the motion direction, while high-frequency edge information perpendicular to the motion direction remains almost undamaged. Applying a uniform inverse offset to all sampling points in the convolutional kernel would result in insufficient inverse pull-back of sampling points along the motion axis, leading to residual tailing in the deblurring. Simultaneously, sampling points perpendicular to the motion axis would be subjected to unnecessary large offsets, introducing lateral artifacts and degraded edge sharpness. Therefore, the technical solution of this application further constructs a normalized motion direction unit vector based on the angle values in the motion direction tensor, and performs an inner product projection and normalization on the coordinate vectors of each sampling point in the convolutional grid with the motion direction unit vector to obtain the direction projection coupling weight matrix. Through the above processing, the spatial distribution differences of the degree of motion blur affecting each sampling point of the convolution kernel can be accurately quantified from the perspective of geometric projection, providing a quantitative physical basis for subsequent implementation of differential offset amplitude modulation for each sampling point.
[0044] More specifically, in a concrete example of this application, the calculation process of the orientation projection coupling weight matrix is as follows. First, based on the pixel-by-pixel motion angle value encoded in the motion direction tensor, cosine and sine functions are calculated for each angle value to construct a normalized motion direction unit vector. This unit vector indicates the precise orientation of motion blur energy propagation at the current pixel position on the two-dimensional plane. Subsequently, for all preset sampling points in the standard deformable convolutional mesh, a fixed mesh coordinate vector relative to the center point is extracted for each sampling point. The dot product is calculated between the mesh coordinate vector and the motion direction unit vector for each sampling point to obtain the scalar projection component of the sampling point on the motion propagation axis. The absolute value of this projection component is taken and normalized based on the maximum projection value among all sampling points to generate an orientation projection coupling weight matrix with values ranging from 0 to 1, as shown below:
[0045]
[0046] in, Indicates the first The directional projection coupling weight scalar corresponding to each sampling point Represents the angle value in the motion direction tensor. The constructed motion direction normalized unit vector, Represents the first convolutional mesh in the standard convolutional grid. A fixed two-dimensional coordinate vector of each sampling point relative to the center point, operator This represents the inner product operation of two-dimensional vectors. This indicates the absolute value operation. This represents the total number of sampling points in the deformable convolution kernel. This represents a very small positive number that prevents the denominator from being zero.
[0047] When the grid coordinate vector of a sampling point is perfectly parallel to the unit vector of the motion direction, the inner product reaches its maximum value, and the weight after normalization approaches one, indicating that the sampling point is located exactly on the axis with the most severe motion blur. When the two are perfectly orthogonal, the inner product is zero, and the weight approaches zero, indicating that the features in the direction of the sampling point are not affected by motion blur. Taking a 3×3 standard convolutional grid as an example, when the motion direction of the assembly line workpiece is horizontal, the grid coordinate vectors of the two sampling points located directly to the left and right of the center point are perfectly parallel to the unit vector of the horizontal motion direction, and their inner product projection reaches its maximum value. After normalization, the weight approaches one, indicating that the feature regions corresponding to these two sampling points carry the most severe blur energy, and subsequent reverse convergence offset with the greatest force needs to be applied. On the other hand, the grid coordinate vectors of the two sampling points located directly above and below the center point are perfectly orthogonal to the unit vector of the horizontal motion direction, the inner product is zero, and the weight approaches zero, indicating that the feature regions corresponding to these two sampling points themselves maintain good edge sharpness, and subsequent perturbation only needs to be applied with a very small amplitude to retain their original high-frequency information. The four sampling points located diagonally receive intermediate weight values between zero and one, reflecting their tilted projection relationship with the motion axis. This weight matrix accurately characterizes the spatial distribution of the degree of motion blur affecting each sampling point of the convolution kernel from the perspective of geometric projection, laying a quantitative foundation for subsequent on-demand convergence.
[0048] In step S322, the directional projection coupling weight matrix is used as a per-sample-point modulation factor to differentially modulate the base offset vector after the motion direction tensor is scaled by the inverse triangular projection and the motion scale tensor, thereby generating an anisotropic inverse offset vector set. It should be noted that, given that the directional projection coupling weight matrix has accurately quantified the geometric coupling degree between each sampling point and the motion propagation axis, if an equal inverse offset is applied to all sampling points, the sampling points arranged along the motion axis will have insufficient inverse pull-back amplitude due to the averaged convergence force, while the sampling points perpendicular to the motion axis will be subjected to unnecessary large offsets, destroying the originally intact lateral high-frequency edge details. Based on this, the technical solution of this application further uses the directional projection coupling weight matrix as a per-sample-point modulation factor to differentially modulate the base offset vector after the motion direction tensor is scaled by the inverse triangular projection and the motion scale tensor, thereby generating an anisotropic inverse offset vector set. Through the above processing, the sampling points arranged along the motion axis can obtain a strong inverse convergence force that matches their diffusion degree, while the sampling points perpendicular to the motion axis are only subject to a very small amplitude perturbation. Thus, the deformable convolution sampling grid is evolved from an isotropic circular convergence mode to an anisotropic elliptical distribution that accurately matches the directional energy diffusion of the real motion blur.
[0049] More specifically, in a concrete example of this application, the generation process of the anisotropic inverse offset vector set is as follows: First, cosine and sine functions are calculated for the angle values in the motion direction tensor and then negatively charged to generate an inverse unit direction component that is strictly opposite to the actual physical motion direction, i.e., an inverse component that differs from the motion direction by π radians. Then, a pixel-by-pixel scalar multiplication is performed on this inverse unit direction component using the motion scale tensor to obtain the basic inverse offset vector. The magnitude of this vector is equal to the complete blur range represented by the motion scale tensor. Finally, the k-th weight in the direction projection coupling weight matrix is used as a modulation factor to perform differential amplitude scaling on the basic inverse offset vector at each sampling point, so that sampling points in different spatial orientations obtain an offset amplitude that matches their degree of blur influence, outputting the anisotropic inverse offset vector set, as shown below:
[0050]
[0051] in, Indicates the first Anisotropic inverse offset vectors corresponding to each sampling point Represents the weight taken from the directional projection coupling weight matrix. The directional projection coupling weights of each sampling point Represents the motion scale tensor. This represents the angle value in the motion direction tensor. and This represents the projection of inverse trigonometric functions, used to generate a projection that differs from the direction of motion. The inverse component of radians.
[0052] when When the offset obtained at the sampling point approaches a certain value, it equals the complete blur range represented by the motion scale tensor, implying a maximum reverse pull-back on the severely diffused features in that direction; when When the offset approaches zero, the offset amplitude is almost zero, and the sampling point remains near its original grid position, thus fully preserving the high-frequency edge details that are undamaged perpendicular to the motion direction. Taking a standard 3×3 convolutional grid as an example, when the motion direction of the assembly line workpiece is horizontal, the two sampling points located directly to the left and right of the center point, due to their directional projection coupling weights approaching one, obtain a reverse horizontal offset equal to the complete blur scale, pulling back the severely diffused horizontal features to their original positions with maximum force. Meanwhile, the two sampling points located directly above and below the center point, due to their weights approaching zero, have almost zero offsets and remain stationary near their original grid coordinates, fully preserving the high-frequency sharpness information of the vertical edges. This differential modulation driven by physical projection relationships allows the deformable convolutional sampling grid to evolve from an isotropic circular convergence mode to an anisotropic elliptical distribution with strong convergence along the motion axis and weak perturbation perpendicular to the motion axis, accurately matching the directional energy diffusion physical model of real motion blur.
[0053] In step S323, the differentiated offset vectors of each sampling point in the anisotropic inverse offset vector set are stacked along the channel dimension and rearranged in the deformable convolution standard offset field format according to the sampling point index order to obtain the inverse spatial offset matrix. It should be noted that, given that differentiated inverse offset vectors have been obtained for each sampling point in the anisotropic inverse offset vector set, but these offset vectors still exist in a discrete, per-sampling-point form, they cannot be directly consumed by the deformable convolution operator and need to be integrated into a unified data structure conforming to the standard two-dimensional offset field format. Based on this, the technical solution of this application further stacks the differentiated offset vectors of each sampling point in the anisotropic inverse offset vector set along the channel dimension and rearranges in the deformable convolution standard offset field format according to the sampling point index order to obtain the inverse spatial offset matrix. Through the above processing, the anisotropic offset information after directional projection coupling modulation can be adapted to a standard input format that the deformable convolution operator can directly superimpose on standard grid coordinates, thereby guiding the actual sampling point distribution of the subsequent anisotropic deformable convolution operator.
[0054] More specifically, in a concrete example of this application, the implementation process of format mapping and offset matrix output is as follows: The differentiated offset vectors of all K sampling points in the anisotropic inverse offset vector set are arranged sequentially from the 1st to the Kth sampling point according to their sampling point indices. The two-dimensional offset vector of each sampling point is then stacked along the channel dimension. Since the offset vector of each sampling point contains two scalar values—a horizontal offset component and a vertical offset component—the stacked K sampling points form a tensor with 2K channels. Subsequently, a format rearrangement operation is performed on this stacked tensor to adapt it to the standard two-dimensional offset field format required by the deformable convolution operator, outputting the inverse spatial offset matrix as follows:
[0055]
[0056] in, This represents the output reverse spatial offset matrix. This represents the differentiated offset vector of each sampling point in the anisotropic inverse offset vector set. Indicates will The two-dimensional offset vectors are stacked and rearranged along the channel dimension. This represents the total number of sampling points for the deformable convolution kernel.
[0057] The rearranged inverse spatial offset matrix is directly superimposed on the standard convolutional grid coordinates, guiding the actual sampling point distribution of subsequent anisotropic deformable convolution operators. Taking a 3×3 standard convolutional grid as an example, K equals 9, and each sampling point has two offset components, horizontal and vertical, so the number of channels in the inverse spatial offset matrix is 18. When the motion direction of the assembly line workpiece is horizontal, the sampling points located directly to the left and right of the center point have values close to the full fuzzy scale in the horizontal offset component, while the vertical offset component is close to zero. The sampling points located directly above and below have offset components close to zero in both directions, while the sampling points at the diagonal positions obtain intermediate offset values between the two. By introducing the directional projection coupling relationship between the sampling point grid coordinates and the motion blur propagation axis, and constructing a sample-point-differentiated anisotropic offset modulation mechanism, the inverse spatial offset matrix can accurately reflect the directional anisotropic physical nature of motion blur. Sampling points arranged along the motion axis achieve strong inverse convergence matching their diffusion level, effectively eliminating residual blur trails caused by insufficient convergence amplitude. Sampling points arranged perpendicular to the motion axis are subject to only minimal perturbation, fully preserving the original undamaged high-frequency edge gradient information in that direction, avoiding lateral artifacts and edge sharpness degradation introduced by uniform offset. Overall, the improved offset matrix-driven deformable convolution sampling grid exhibits an anisotropic elliptical distribution that precisely matches the real motion blur energy diffusion model. Without increasing additional network parameters or inference computation, it simultaneously improves the thoroughness of deblurring in the motion direction and the edge fidelity in the orthogonal direction, thereby enabling the subsequent bounding box regression of the detector head to obtain higher-precision positioning references in both orthogonal dimensions.
[0058] In step S33, based on the inverse spatial offset matrix, anisotropic inverse convergence resampling and edge feature weighted sharpening are performed on the shallow defect semantic feature tensor to obtain an inverse aggregated sharpened feature tensor. It should be noted that since the inverse spatial offset matrix has already encoded the differentiated anisotropic inverse offset information of each sampling point, and the high-frequency energy of the defect edges in the shallow defect semantic feature tensor is still diffuse along the direction of motion, this offset matrix needs to be used to drive the deformable convolution operator to complete the actual inverse convergence sampling and aggregation operation in the feature space. Based on this, the technical solution of this application further performs anisotropic inverse convergence resampling and edge feature weighted sharpening on the shallow defect semantic feature tensor based on the inverse spatial offset matrix to obtain an inverse aggregated sharpened feature tensor. Through the above processing, the diffused defect edge energy can be inversely aggregated back to its original spatial position in the high-dimensional feature space according to the anisotropic convergence path, reconstructing a feature expression with sharp edges and clear contrast.
[0059] More specifically, in a concrete example of this application, the implementation process of anisotropic inverse convergence resampling and edge feature weighted sharpening is as follows. First, the inverse spatial offset matrix is superimposed on the fixed preset coordinates of each sampling point in the standard convolutional grid to calculate the actual sampling landing coordinates of each sampling point in the anisotropic deformable convolution operator. Since the offset of each sampling point exhibits a differentiated distribution after being modulated by the directional projection coupling weight, the sampling points arranged along the motion axis are pulled to the opposite position away from the center to recover diffuse energy, while the sampling points perpendicular to the motion axis only produce a minimal displacement to maintain the integrity of the in-situ features. The overall sampling grid exhibits an anisotropic elliptical distribution. Subsequently, since the actual sampling landing coordinates after superimposition and offset fall at non-integer pixel positions, the feature values at the corresponding positions in the shallow defect semantic feature tensor are extracted with sub-pixel precision using a bilinear interpolation algorithm. The accurate interpolation response is calculated by weighting the feature values at the four nearest neighbor integer coordinates of the diffuse region. Finally, combining the modulation scalar and convolution weight matrix automatically learned by the network during training, the feature responses extracted from each sampling point are weighted, summed, and aggregated to obtain the inverse aggregated sharpened feature tensor. The computational logic of this aggregation process is as follows:
[0060]
[0061] in, This indicates the location of the inverse aggregation sharpening feature tensor in the central space. The output feature vector at that location, This represents the total number of sampling points for anisotropic deformable convolution kernels. Indicates the first The learnable convolutional weight matrix corresponding to each sampling point This represents the feature response value at the corresponding coordinate in the semantic feature tensor of shallow defects. Represents the first in a standard convolutional mesh Each sampling point has a fixed preset coordinate relative to the center point. Represents the first element resolved from the reverse space offset matrix. Anisotropic dynamic offset of each sampling point The learnable attention modulation scalar is used to control the aggregation contribution of features at this bias position to the center point. The introduction of the modulation scalar gives higher aggregation weights to the diffuse features acquired along the motion axis, further enhancing the edge sharpening effect of inverse convergence. After the above weighted aggregation, the high-frequency gradient information of the defect edges in the inverse aggregation sharpening feature tensor has been restored from a diffuse state to a spatially concentrated and sharply contrasting distribution. At the same time, the edge details perpendicular to the motion direction are preserved in situ due to the minimal offset of the corresponding sampling points. The overall feature representation has clear boundary localization benchmarks in both orthogonal dimensions.
[0062] Specifically, in step S4, the inverse aggregated sharpening feature tensor is subjected to cross-scale feature alignment and multi-scale semantic fusion through the top-down semantic transfer path and bottom-up spatial localization path of the feature pyramid network to obtain a multi-scale defect feature pyramid. It should be noted that, due to the wide range of physical dimensions of defects on industrial production lines—a tiny crack may occupy only a few pixels while a long scratch may span tens of pixels—a single-scale inverse aggregated sharpening feature tensor cannot simultaneously meet the detection requirements of defects of different physical sizes. Therefore, the technical solution of this application further uses the top-down semantic transfer path and bottom-up spatial localization path of the feature pyramid network to perform cross-scale feature alignment and multi-scale semantic fusion on the inverse aggregated sharpening feature tensor to obtain a multi-scale defect feature pyramid. Through the above processing, a multi-scale feature representation that simultaneously contains high-resolution precise spatial localization information and low-resolution global abstract semantic information can be constructed, enabling the subsequent detection head to effectively detect defects of different physical sizes.
[0063] More specifically, in a concrete example of this application, the inverse aggregated sharpening feature tensor is first subjected to multi-level spatial resolution downgrading and channel normalization compression to obtain a hierarchical initial feature set. The inverse aggregated sharpening feature tensor is fed in parallel into multiple branch paths. The first branch maintains the original spatial resolution unchanged, the second branch reduces the spatial resolution to half of the original using a max pooling operator with a stride of 2, and the third branch reduces the spatial resolution to one-quarter of the original using a max pooling operator with a stride of 4. A 1×1 linear projection convolution is applied to the feature maps of different resolutions output by each branch to normalize and activate the number of channels, resulting in a hierarchical initial feature set containing three scale dimensions: shallow high resolution, mid-level, and deep low resolution.
[0064] Subsequently, the hierarchical initial feature set is fused stepwise from deep global semantics to shallow features to obtain a semantically enhanced intermediate feature set. Starting with the deep features in the hierarchical initial feature set that have the lowest spatial resolution but the highest level of semantic abstraction, a nearest neighbor interpolation algorithm with a scaling factor of 2 is used to spatially upsample them, aligning their resolution with that of the adjacent mid-level features. The upsampled deep features are then added element-wise with the mid-level features, injecting the deep global category semantic information into the mid-level features. Following the same upsampling and element-wise addition logic, the fused mid-level features are then passed to and fused with the shallow high-resolution features, resulting in a semantically enhanced intermediate feature set where each level is rich in global abstract semantics.
[0065] Finally, bottom-up spatial localization feature downsampling and cross-scale aggregation are performed on the semantically enhanced intermediate feature set to obtain a multi-scale defect feature pyramid. Starting from the shallowest features with the highest resolution in the semantically enhanced intermediate feature set, spatial compression is performed using a 3×3 downsampling convolution with a stride of 2. The compressed shallow features are then concatenated with the original mid-level features of the same level along the channel dimension, followed by channel dimensionality reduction and feature alignment using a 1×1 convolution. This downsampling and concatenation process is repeated to advance to deeper levels, allowing precise spatial geometric localization information to be progressively transferred from the shallow to the deep layers, ultimately outputting the multi-scale defect feature pyramid. The features at each level of this pyramid simultaneously possess global semantic information from the top-down path and precise spatial localization information from the bottom-up path, providing sufficient feature support for subsequent decoupled detection heads targeting industrial defects of different physical sizes.
[0066] Specifically, in step S5, the defect category and location are predicted based on the decoupled detection head on the multi-scale defect feature pyramid to output a defect localization detection result set. It should be noted that, given that the multi-scale defect feature pyramid has already fused edge features sharpened by inverse aggregation with multi-scale semantic information, but defect classification and bounding box localization belong to two different tasks with different optimization objectives—the former focusing on semantic discrimination and the latter on geometric coordinate regression—sharing the same feature mapping path will lead to gradient conflicts. Furthermore, in motion-blurred scenarios, the confidence of boundary localization is subject to jitter due to residual diffusion, necessitating the introduction of an uncertainty perception mechanism to dynamically penalize high-jitter prediction boxes. Based on this, the technical solution of this application further predicts the defect category and location based on the decoupled detection head on the multi-scale defect feature pyramid to output a defect localization detection result set. Through the above processing, gradient coupling interference between classification and regression tasks can be eliminated. Simultaneously, through the synchronous calculation of boundary uncertainty variance and the dynamic penalty filtering mechanism, the jitter in prediction box localization caused by motion blur residue is effectively suppressed, outputting defect detection results with both high classification accuracy and high localization stability.
[0067] Figure 6This is a flowchart illustrating a convolutional neural network-based industrial defect detection method according to an embodiment of this application, which predicts the defect category and location based on a multi-scale defect feature pyramid using a decoupled detection head to output a defect localization detection result set. Figure 6 As shown, step S5 includes: S51, using the classification detection head and regression detection head, performing feature distribution and category probability distribution extraction on the multi-scale defect feature pyramid to obtain the defect category probability distribution tensor and decoupled regression intermediate features; S52, simultaneously calculating the bounding box geometric coordinates and localization uncertainty variance on the decoupled regression intermediate features to obtain the bounding box prediction coordinate tensor and boundary uncertainty variance tensor; S53, based on the boundary uncertainty variance tensor, combining the defect category probability distribution tensor and the bounding box prediction coordinate tensor, performing intersection-union ratio calculation and uncertainty weighted nonmaximum suppression filtering to output the defect localization detection result set.
[0068] In step S51, the classification and regression detection heads are used to perform feature distribution and class probability distribution extraction on the multi-scale defect feature pyramid to obtain the defect class probability distribution tensor and decoupled regression intermediate features. It should be noted that since the defect classification task and the bounding box regression task have fundamentally different optimization objectives—the classification task focuses on semantic discrimination while the regression task focuses on geometric coordinate accuracy—sharing the same feature mapping path will lead to gradient conflicts and mutual interference. Therefore, the technical solution of this application further uses the classification and regression detection heads to perform feature distribution and class probability distribution extraction on the multi-scale defect feature pyramid to obtain the defect class probability distribution tensor and decoupled regression intermediate features. Through the above processing, the gradient coupling effect between the two tasks can be severed, allowing the classification branch to focus on defect semantic discrimination while the regression branch focuses on optimizing geometric positioning accuracy.
[0069] More specifically, in a concrete example of this application, the implementation process of feature distribution and category probability distribution extraction is as follows: The multi-scale defect feature pyramid is distributed to the structurally independent classification and regression detection heads according to each level. Within the classification detection head, multiple sets of 3×3 convolution operators are sequentially applied to the input features at each level for local spatial dimensionality reduction and receptive field integration, gradually compressing high-dimensional features into a compact expression oriented towards category discrimination. Finally, a 1×1 convolution layer maps the number of channels to the preset number of industrial defect categories, and a Sigmoid activation function is applied to the output to map each channel value to an independent confidence score between zero and one, obtaining a defect category probability distribution tensor. Each spatial anchor point in this tensor corresponds to the classification probability value of each type of industrial defect. Simultaneously, within the regression detection head, multiple sets of 3×3 convolution operators are also applied to the input features at each level for spatial feature integration and dimensionality reduction, extracting decoupled regression intermediate features specifically for subsequent bounding box coordinate parameter and uncertainty variance calculation. The two paths are completely independent in terms of network structure and do not share any convolutional layer parameters, ensuring that the classification gradient and regression gradient do not interfere with each other during backpropagation.
[0070] In step S52, the bounding box geometric coordinates and positioning uncertainty variance are simultaneously calculated on the decoupled regression intermediate features to obtain the bounding box predicted coordinate tensor and the boundary uncertainty variance tensor. It should be noted that even after feature-level inverse aggregation deblurring in high-speed pipeline scenarios, some severely blurred areas may still have residual diffusion, causing positioning jitter in the bounding box coordinate regression results in these areas. Traditional regression heads only output deterministic coordinate values and cannot quantify this jitter risk. Therefore, the technical solution of this application further performs simultaneous calculation of the bounding box geometric coordinates and positioning uncertainty variance on the decoupled regression intermediate features to obtain the bounding box predicted coordinate tensor and the boundary uncertainty variance tensor. Through the above processing, a confidence quantification index can be provided for each coordinate component while outputting geometric positioning parameters, providing a numerical basis for subsequent dynamic penalty filtering of high-jitter prediction boxes.
[0071] More specifically, in a concrete example of this application, the implementation process of simultaneously calculating bounding box coordinates and uncertainty variance is as follows: Decoupled regression intermediate features are simultaneously fed into parallel coordinate parameter regression layers and uncertainty estimation layers. In the coordinate parameter regression layer, a 1×1 convolution kernel maps the number of channels of the decoupled regression intermediate features to 4, corresponding to the horizontal and vertical offsets of each anchor point relative to the center of the feature map grid, as well as the absolute width and height of the detection box, directly outputting the bounding box prediction coordinate tensor. In the uncertainty estimation layer, another independent set of 1×1 convolution kernels also maps the number of channels to 4, outputting a logarithmic variance parameter corresponding one-to-one with the four coordinate components. This parameter, in logarithmic space, characterizes the degree of positional confidence fluctuation caused by motion blur residues in each coordinate component. Subsequently, a natural exponential function is applied to this logarithmic variance parameter to transform it into a strictly non-negative variance value, obtaining the boundary uncertainty variance tensor. The calculation logic of this variance calculation is expressed as follows:
[0072]
[0073] in, This represents the boundary uncertainty variance tensor of the output. This represents the intermediate features of the decoupled regression input. This represents the two-dimensional convolution operator. This represents the 1×1 convolution kernel weight matrix used for estimating the log-variance of boundary uncertainties. This indicates the corresponding bias term. The natural exponential function is used to ensure the strict non-negativity of the variance output value. A larger variance value indicates that the regression result for that coordinate component is more severely affected by fuzzy residuals and has a higher risk of positioning jitter; a smaller variance value indicates that the regression result for that coordinate component has higher positioning reliability. The two regression paths share decoupled regression intermediate features as input but each has independent convolution parameters, ensuring the independence of coordinate regression and variance estimation in the gradient optimization process.
[0074] In step S53, based on the boundary uncertainty variance tensor, the intersection-union ratio (IU / R) is calculated and uncertainty-weighted non-maximum suppression (NMS) is applied to filter out redundant boxes using the defect category probability distribution tensor and the bounding box prediction coordinate tensor to output a defect localization detection result set. It should be noted that in high-speed pipeline scenarios, multiple spatially overlapping candidate predicted boxes often appear for the same defect region. Some of these predicted boxes have high boundary uncertainty variance due to motion blur residue. If traditional NMS is used to filter out boxes based solely on classification confidence and IU / R, it cannot distinguish between predicted boxes with high localization confidence and those with high jitter risk. Therefore, the technical solution of this application further calculates the IU / R and applies uncertainty-weighted NMS to filter out redundant boxes using the boundary uncertainty variance tensor, combining the defect category probability distribution tensor and the bounding box prediction coordinate tensor to output a defect localization detection result set. Through the above processing, predicted boxes with high localization uncertainty can be dynamically penalized during redundant box filtering, while retaining detection results with stable localization and accurate classification.
[0075] More specifically, in a concrete example of this application, the implementation process of uncertainty-weighted nonmaximum suppression is as follows. First, the defect category probability distribution tensor, the bounding box prediction coordinate tensor, and the boundary uncertainty variance tensor are dimensionally aligned and bound along their corresponding spatial anchor points, reorganizing them into a set of candidate boxes with all attributes, including category scores, four-dimensional coordinate parameters, and four-dimensional variance parameters. Then, for any given defect category, the predicted box with the highest current category probability is selected from the candidate box set as the best reference box, and the intersection-union ratio (IUR) between the remaining surrounding candidate boxes and this best reference box is calculated. In traditional nonmaximum suppression, candidate boxes with an IUR exceeding a preset threshold are directly eliminated. This scheme introduces a dynamic penalty mechanism using the boundary uncertainty variance tensor, substituting the IUR and the variance value of the corresponding candidate box into a Gaussian decay function to dynamically update the final scores of the surrounding candidate boxes. The calculation logic of this dynamic penalty is expressed as follows:
[0076]
[0077] in, This represents the final score of the candidate boxes after dynamic penalty. This represents the original confidence score in the probability distribution tensor of the defect category corresponding to the candidate box. This represents the intersection-union ratio (CIRR) function used to calculate the intersection-union ratio of two predicted bounding boxes. This represents the bounding box prediction coordinate tensor corresponding to the surrounding candidate boxes. This represents the best reference bounding box with the highest confidence level in the current category. This represents the boundary uncertainty variance tensor corresponding to the candidate box. This represents a very small positive number to prevent the denominator from being zero. When the boundary uncertainty variance of a candidate box is large, increasing the denominator causes the Gaussian decay factor to approach one, weakening the penalty effect. However, since the candidate box itself has low localization confidence, its original confidence score has already been suppressed during training, resulting in a still low overall score. When the intersection-union ratio of a candidate box with the best reference box is high and its own variance is small, the Gaussian decay factor has a stronger penalty, effectively suppressing highly overlapping redundant boxes. Finally, a filtering threshold is set based on the final score after penalty adjustment. Candidate boxes with scores below the threshold are removed, and predicted boxes with scores above the threshold are retained as the final output defect localization detection result set. Each retained detection box in this result set contains a defect category label, four-dimensional bounding box coordinates, and the corresponding classification confidence score, marking the completion of the closed loop of the entire high-speed pipeline defect detection process.
[0078] Specifically, in a concrete example of this application, the network training process of the above-mentioned industrial defect detection method based on convolutional neural networks is as follows. The training dataset consists of labeled grayscale images acquired by an industrial linear scan camera on a high-speed production line. The annotation information of each image includes a defect category label and the corresponding ground truth coordinates of the bounding box. The total loss function during the training phase consists of a weighted sum of three parts: classification loss, bounding box regression loss, and uncertainty regularization loss. The classification loss uses the Focal Loss function, which dynamically reduces the loss weight of easily classified samples through a modulation factor to alleviate the problem of extreme imbalance between positive and negative samples in industrial scenarios. The bounding box regression loss uses the uncertainty-weighted Smooth L1 loss, which uses the inverse of the boundary uncertainty variance tensor as the adaptive weight of the regression loss of each coordinate component, so that the network automatically reduces the regression gradient contribution of high uncertainty regions during training and avoids the interference of noise labels in severely blurred regions on the overall regression accuracy. The uncertainty regularization loss applies a logarithmic penalty term to the boundary uncertainty variance tensor to prevent the network from circumventing the regression loss by infinitely increasing the variance. The network uses the Adam optimizer for parameter updates, with an initial learning rate set to 1×10⁻ 4 A cosine annealing strategy is employed to gradually decay the learning rate during training. The training batch size is set to 16, and the total training epochs are 300. An early stopping mechanism is triggered when the mAP metric on the validation set stops improving for 20 consecutive epochs. The motion vector perception branch and the backbone detection network adopt an end-to-end joint training strategy, eliminating the need for staged pre-training.
[0079] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. An industrial defect detection method based on convolutional neural networks, characterized in that, include: Step 1: Extract shallow fuzzy defect features from the motion-blurred grayscale image of the production line acquired by an industrial line scan camera on a high-speed production line to obtain the shallow defect semantic feature tensor. Step 2: Perform defect kinematic feature solving based on local perception branch on the shallow defect semantic feature tensor to obtain the local motion vector field map; Step 3: Based on the motion direction angle and fuzzy scale parameters encoded in the local motion vector field map, perform feature-level inverse aggregation defuzzification on the shallow defect semantic feature tensor based on motion dynamics constraints to obtain the inverse aggregation sharpened feature tensor. Step 4: Through the top-down semantic transmission path and bottom-up spatial localization path of the feature pyramid network, cross-scale feature alignment and multi-scale semantic fusion are performed on the inverse aggregated sharpening feature tensor to obtain a multi-scale defect feature pyramid. Step 5: Based on the decoupled detection head, perform defect category and location prediction on the multi-scale defect feature pyramid to output a defect localization detection result set.
2. The industrial defect detection method based on convolutional neural networks according to claim 1, characterized in that, The motion blur grayscale image of the pipeline includes the distribution of diffuse pixels in the direction of motion and the grayscale gradient distribution in the defect area.
3. The industrial defect detection method based on convolutional neural networks according to claim 2, characterized in that, Step one includes: Step 1.1: Based on a 7×7 large receptive field two-dimensional convolution kernel with a stride of 2, perform sliding convolution and nonlinear activation mapping on the pipeline motion blur grayscale image to obtain the initial high-frequency response feature map. Step 1.2: Using a residual downsampling module containing a 3×3 convolutional main branch and a 1×1 identity mapping branch, the initial high-frequency response feature map is spatially downsampled layer by layer and local topological features are extracted to obtain the intermediate topological aggregated feature tensor. Step 1.3: Using a 1×1 pointwise convolution operator, multi-channel feature fusion and shallow semantic reconstruction are performed on the intermediate topological aggregate feature tensor to obtain the shallow defect semantic feature tensor.
4. The industrial defect detection method based on convolutional neural networks according to claim 1, characterized in that, Step two includes: Step 2.1: Perform motion fuzzy prior enhancement based on frequency domain transformation on the shallow defect semantic feature tensor to obtain the frequency domain enhanced feature tensor; Step 2.2: Using two sets of parallel 1×1 convolution kernels, perform pixel-wise dual-track regression of motion direction and blur scale on the frequency domain enhancement feature tensor to obtain the motion parameter mapping tensor; Step 2.3: Using a 3×3 depth-separable convolution operator, the motion parameter mapping tensor is spatially smoothed to ensure consistent motion parameters between adjacent pixels to obtain a local motion vector field map.
5. The industrial defect detection method based on convolutional neural networks according to claim 1, characterized in that, Step three includes: Step 3.1: Perform feature slicing and decoupling on the local motion vector field map along the channel dimension to obtain the motion direction tensor and motion scale tensor; Step 3.2: Determine the reverse spatial offset matrix based on the motion direction tensor and the motion scale tensor; Step 3.3: Based on the inverse spatial offset matrix, anisotropic inverse convergence resampling and edge feature weighted sharpening are performed on the shallow defect semantic feature tensor to obtain the inverse aggregated sharpened feature tensor.
6. The industrial defect detection method based on convolutional neural networks according to claim 1, characterized in that, Step four includes: Step 4.1: Perform multi-level spatial resolution downgrading and channel normalization compression on the inverse aggregated sharpening feature tensor to obtain a hierarchical initial feature set; Step 4.2: Perform deep global semantic injection and fusion into the shallow layer of the hierarchical initial feature set to obtain a semantically enhanced intermediate feature set; Step 4.3: Perform bottom-up spatial localization feature downsampling and cross-scale aggregation on the semantically enhanced intermediate feature set to obtain a multi-scale defect feature pyramid.
7. The industrial defect detection method based on convolutional neural networks according to claim 1, characterized in that, Step five includes: Step 5.1: Using the classification detection head and the regression detection head, feature distribution and category probability distribution extraction are performed on the multi-scale defect feature pyramid to obtain the defect category probability distribution tensor and decoupled regression intermediate features. Step 5.2: Simultaneously calculate the bounding box geometric coordinates and the localization uncertainty variance of the decoupled regression intermediate features to obtain the bounding box predicted coordinate tensor and the boundary uncertainty variance tensor. Step 5.3: Based on the boundary uncertainty variance tensor, the intersection-union ratio (IUU) is calculated and uncertainty-weighted nonmaximum suppression filtering is performed by combining the defect category probability distribution tensor and the bounding box prediction coordinate tensor to output the defect localization detection result set.
8. The industrial defect detection method based on convolutional neural networks according to claim 5, characterized in that, Step 3.2 includes: Step 3.2.1: Based on the angle values in the motion direction tensor, construct a normalized motion direction unit vector, and perform inner product projection and normalization on the coordinate vectors of each sampling point of the convolutional grid and the motion direction unit vector to obtain the direction projection coupling weight matrix. Step 3.2.2: Using the direction projection coupling weight matrix as the modulation factor for each sampling point, differential amplitude modulation is performed on the basic offset vector after the motion direction tensor is scaled by the inverse triangular projection and the motion scale tensor to generate an anisotropic inverse offset vector set. Step 3.2.3: Stack the differential offset vectors of each sampling point in the anisotropic inverse offset vector set according to the sampling point index order, and rearrange them in the channel dimension and deformable convolution standard offset field format to obtain the inverse space offset matrix.