A steel plate surface defect detection system based on space-time mutual attention and sparse space-time perception attention
By using a detection system based on spatiotemporal mutual attention and sparse spatiotemporal perception attention, the problems of adaptive and dynamic correlation modeling in steel plate surface defect detection are solved, achieving efficient and accurate defect identification and adapting to the detection needs of complex industrial scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies for detecting defects on steel plate surfaces lack adaptability, struggle to effectively characterize multi-scale and fuzzy edge features, lack dynamic correlation modeling capabilities in spatial and temporal dimensions, and have insufficient model generalization performance and stability, making it particularly difficult to identify diverse defects in complex industrial scenarios.
A detection system based on spatiotemporal mutual attention and sparse spatiotemporal awareness attention is adopted, including steel plate image acquisition, shallow feature extraction, hybrid deformable convolution, multi-scale feature stitching, spatiotemporal collaborative feature fusion, cross attention and Informer decoder. The training process is optimized by combining spatiotemporal collaborative feature fusion and sparse spatiotemporal awareness self-attention mechanism with dynamic weighted total loss function.
It improves the accuracy and stability of steel plate surface defect detection, enhances the adaptability to complex defect morphologies and detection efficiency, reduces computational complexity, improves detection accuracy and generalization performance, and adapts to application needs in different production environments.
Smart Images

Figure CN121213558B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of steel plate surface defect detection technology, specifically to a steel plate surface defect detection system based on spatiotemporal mutual attention and sparse spatiotemporal perception attention. Background Technology
[0002] In the steel production process, the surface quality of steel plates directly affects the product's performance and market competitiveness. Therefore, the detection and identification of surface defects in steel plates is of great significance. Traditional detection methods mainly rely on manual visual inspection or algorithms based on simple image processing. However, these methods are inefficient in large-scale production environments and are easily affected by the subjective factors of operators, making it difficult to guarantee stability and consistency.
[0003] With the development of computer vision and deep learning technologies, convolutional neural networks (CNNs) have been increasingly applied to steel plate surface defect detection, enabling a degree of automated detection. However, existing CNN-based detection methods still have significant limitations when facing complex industrial scenarios: conventional convolutional structures lack sufficient adaptability for diverse defect morphologies such as cracks, oxide spots, and inclusions, making it difficult to effectively characterize multi-scale and blurred edge features; simultaneously, traditional networks have limited ability to model the dynamic correlations of steel plate surface defects in spatial and temporal dimensions, making it difficult to capture the evolutionary patterns of defects over time or production status. Furthermore, industrial data often features sparse defect samples, class imbalance, and significant differences between different production lines, resulting in insufficient generalization performance and detection stability. Although attention mechanism models have shown good potential in temporal feature modeling in recent years, their research and application in industrial surface defect detection remain insufficient. Summary of the Invention
[0004] To address the problems existing in the prior art, this invention provides a steel plate surface defect detection system based on spatiotemporal mutual attention and sparse spatiotemporal perception attention, which significantly improves the accuracy, stability and generalization ability of steel plate surface defect detection in complex production environments, and provides a new technical path for industrial quality inspection.
[0005] To achieve the above technical objectives, the present invention adopts the following technical solution:
[0006] A steel plate surface defect detection system based on spatiotemporal mutual attention and sparse spatiotemporal perception attention includes: a steel plate image acquisition module, a shallow feature extraction module, a hybrid deformable convolution module, a multi-scale feature stitching module, a spatiotemporal collaborative feature fusion module, a cross-attention module, a feature stitching module, and an Informer decoder;
[0007] The steel plate image acquisition module is used to acquire images of the steel plate surface.
[0008] The shallow feature extraction module is used to extract shallow semantic features of the steel plate surface from the steel plate surface image;
[0009] The hybrid deformable convolution module is used to perform multi-scale feature extraction on the extracted shallow semantic features of the steel plate surface;
[0010] The multi-scale feature stitching module is used to stitch together multi-scale features;
[0011] The spatiotemporal collaborative feature fusion module is used to extract the spatial-temporal collaborative features of the steel plate surface from the spliced multi-scale features using a spatiotemporal mutual attention mechanism, and to fuse the spatial-temporal collaborative features of the steel plate surface with the spliced multi-scale features;
[0012] The cross-attention module is used to extract the features at the intersection of high-resolution texture features and low-resolution semantic features on the surface of the steel plate using a cross-attention mechanism;
[0013] The feature splicing module is used to splice fused features and intersecting features;
[0014] The Informer decoder is used to extract sparse focused features of the steel plate surface from the spliced features using a sparse spatiotemporal awareness self-attention mechanism, and to predict the type of defects on the steel plate surface based on the sparse focused features of the steel plate surface.
[0015] Furthermore, the shallow feature extraction module includes a block embedding unit, a layer normalization unit, and a layer normalization unit connected in sequence. Convolutional unit;
[0016] The block embedding unit is used to divide the acquired steel plate surface image into non-overlapping image patches, and embed all image patches into a fixed-dimensional feature vector through linear mapping;
[0017] The layer normalization unit is used to normalize the feature vector;
[0018] The Convolutional units are used to extract shallow semantic features of the steel plate surface from normalized feature vectors.
[0019] Furthermore, the hybrid deformable convolution module includes: Deformable convolutional unit Deformable convolutional units and Deformable convolutional unit;
[0020] The Deformable convolutional units utilize layer scaling to perform channel-by-channel scaling on the extracted shallow semantic features of the steel plate surface, and then utilize... Deformable convolutional layers extract local detail enhancement features from the surface of steel plates;
[0021] The Deformable convolutional units utilize adaptive activation functions to perform nonlinear transformations on local detail enhancement features, and then utilize... Deformable convolutional layers extract mesoscale features from the surface of steel plates;
[0022] The Deformable convolutional units normalize mesoscale features using layer normalization, and then utilize... Deformable convolutional layers extract global features from the surface of steel plates.
[0023] Furthermore, the process by which the cross-attention module extracts the intersection of high-resolution texture features and low-resolution semantic features on the steel plate surface using a cross-attention mechanism is as follows:
[0024] i: Upsampling is used to obtain high-resolution texture features of the steel plate surface by enhancing local details. The global features of the steel plate surface are downsampled to obtain low-resolution semantic features of the steel plate surface. ;
[0025] ii: High-resolution texture features on the steel plate surface Mapping to query matrix Bond matrix Low-resolution semantic features of steel plate surface Mapping to value matrix The cross-attention mechanism is used to extract the features at the intersection of high-resolution texture features and low-resolution semantic features on the surface of steel plates. :
[0026]
[0027] in, These are the first learnable weight matrix, the second learnable weight matrix, and the third learnable weight matrix, respectively. is the scaling factor for the key matrix. This indicates the transpose operation.
[0028] Furthermore, the process by which the spatiotemporal collaborative feature fusion module extracts the spatial-temporal collaborative features of the steel plate surface using a spatiotemporal mutual attention mechanism is as follows:
[0029] i: Perform temporal modeling on the spliced multi-scale features to extract the temporal context features of the steel plate surface. and the time context features of the steel plate surface Mapping to the temporal feature key matrix With time series eigenvalue matrix ;
[0030] ii: Spatial modeling of the spliced multi-scale features to extract spatial features from the steel plate surface. and the spatial features of the steel plate surface Mapping to spatial feature query matrix ;
[0031] iii: Transform the temporal feature key matrix Temporal eigenvalue matrix Spatial feature query matrix The spatial-temporal synergistic features of the steel plate surface are integrated:
[0032]
[0033] in, The scaling factor for the time series feature key matrix. For normalized exponential functions, This indicates the transpose operation.
[0034] Furthermore, the process of fusing the spatial-temporal synergistic features of the steel plate surface with the multi-scale features of the splicing is as follows:
[0035]
[0036] in, Indicates the characteristics of fusion, This represents the multi-scale features of the splicing. express The learnable balance coefficient, express The learnable balance coefficient.
[0037] Furthermore, the Informer decoder includes: a sparse spatiotemporal awareness self-attention mechanism module, a first-layer normalization unit, a feedforward neural network, and a second-layer normalization unit;
[0038] The sparse spatiotemporal perception self-attention mechanism module is used to extract sparse focused features on the surface of the steel plate from the features spliced by the feature splicing module;
[0039] The first normalization unit is used to normalize the sparse focusing features on the surface of the steel plate;
[0040] The feedforward neural network is used to perform a nonlinear transformation on the normalized sparse focused features;
[0041] The second normalization unit is used to normalize the features of the nonlinear transformation and predict the probability of the type of defect on the steel plate surface.
[0042] Furthermore, the process of extracting the sparse focusing features of the steel plate surface is as follows:
[0043] i: Unify the dimensions of the features spliced by the feature splicing module, flatten them into a spatiotemporal token sequence, and map the spatiotemporal token sequence into a spatiotemporal token sequence value matrix;
[0044] ii: For each spatiotemporal token in the spatiotemporal token sequence, the coordinate information of the query position corresponding to the spatiotemporal token and the time step to which it belongs are included, and the spatiotemporal token is mapped to a spatiotemporal token query matrix and a spatiotemporal token key matrix;
[0045] iii: For any two query positions in the spatiotemporal token sequence, introduce a spatiotemporal sparse mask threshold and calculate the sparse attention weights of the query positions:
[0046]
[0047] in, Indicates the query position in the spacetime token sequence and query location Sparse attention weights, Indicates the query position The spatiotemporal token query matrix Indicates the query position The spacetime token key matrix, This indicates the transpose operation. Indicates the query position The scaling factor of the spacetime token key matrix. Indicates the query position Coordinate information, Indicates the query position Coordinate information, Indicates the query position The time step to which it belongs Indicates the query position The time step to which it belongs This represents the spatial sparse mask threshold in the spatiotemporal sparse mask thresholding. This represents the temporal sparse mask threshold in the spatiotemporal sparse mask threshold;
[0048] iv: Introduce sparse attention weights corresponding to the query position into the spatiotemporal token sequence value matrix to obtain sparse focusing features.
[0049] Furthermore, using historically collected steel plate surface images as input, and the types of steel plate surface defects on the historically collected steel plate surface images as labels, the steel plate surface defect detection system is trained until the total loss function, which integrates normal region loss, defect region loss, and neighborhood adaptive loss, converges, thus completing the training of the steel plate surface defect detection system.
[0050] Furthermore, the total loss function Represented as:
[0051]
[0052] in, The loss function represents the normal area of the steel plate surface. , This represents the number of samples from the normal area of the steel plate surface. for index, For the first Predicted values for samples from normal areas on the surface of a steel plate. For the first Labels on samples from normal areas of a steel plate surface; The loss function for the defect region on the steel plate surface. , This represents the total number of surface defect types on the steel plate. express index, For the first Number of samples of surface defects in steel plates for index, For the first The severity of surface defects in steel plates. For the first The first sample of the surface defect area of the steel plate Predicted values of surface defects in steel plates For the first The sample from the surface defect area of the steel plate belongs to the first... Labeling of surface defects on steel plates for The adjustment coefficient; This represents the domain-adaptive loss function. , This represents the mean of all predicted values for samples from normal areas and defective areas on the steel plate surface. This represents the covariance of all predicted values for samples from normal areas and samples from defective areas on the steel plate surface. This represents the mean of all predicted values of a pseudo-sample constructed using Gaussian blurring with the domain enhancement operator, combining samples from normal areas and defective areas of the steel plate surface. This represents the covariance of all predicted values of a pseudo-sample constructed using Gaussian blurring with the domain enhancement operator, by combining samples from normal areas and defective areas of the steel plate surface. Denotes the Frobenius norm. for The adjustment coefficient.
[0053] Compared with the prior art, the present invention has the following beneficial effects:
[0054] (1) The steel plate surface defect detection system based on spatiotemporal mutual attention and sparse spatiotemporal perception attention designed in this invention has a spatiotemporal collaborative feature fusion module, which realizes the bidirectional fusion of spatial features and temporal context features of the steel plate surface, and can effectively capture the spatiotemporal variation law of steel plate surface defects, and improve the adaptability and expressive ability of the steel plate surface defect detection system to complex defect morphology.
[0055] (2) The steel plate surface defect detection system based on spatiotemporal mutual attention and sparse spatiotemporal perception attention of the present invention introduces a sparse spatiotemporal perception self-attention mechanism in the Informer decoder. Based on sparse prior, it focuses on high confidence areas on the steel plate surface, thereby significantly reducing computational complexity and improving detection efficiency while maintaining detection accuracy.
[0056] (3) The steel plate surface defect detection system based on spatiotemporal mutual attention and sparse spatiotemporal perception attention of the present invention introduces a dynamic weighted total loss function that integrates normal region loss, defect region loss and domain adaptive loss during the training process. The normal region loss is used to maintain the stable learning of the overall texture and non-defect region features of the steel plate surface by the steel plate surface defect detection system, and avoid the feature distribution shift caused by excessive attention to local anomalies in the early stage of training. The defect region loss introduces the dynamic weight and severity coefficient of the defect category, so that the steel plate surface defect detection system can gradually enhance the learning of sparse, subtle and complex defect samples during the training process, and improve the detection sensitivity and recognition accuracy. The domain adaptive loss compares the feature distribution of the original sample and the sample enhanced by Gaussian blur, and constrains the consistency of the mean and covariance of the feature space, so that the steel plate surface defect detection system maintains a relatively stable feature expression under different lighting, texture and production conditions. Through the synergistic effect of these three factors, the steel plate surface defect detection system can take into account the stability of the global structure, the ability to identify local defects, and the adaptability to different environments during the feature learning process. This enables adaptive optimization of the learning weights of each region during the training process, thereby improving the generalization performance and cross-domain adaptability of the steel plate surface defect detection system. Attached Figure Description
[0057] Figure 1 This is a schematic diagram of the steel plate surface defect detection system based on spatiotemporal mutual attention and sparse spatiotemporal sensing attention according to the present invention.
[0058] Figure 2 This is a comparison chart showing the accuracy of the present invention in detecting defects on the surface of steel plates with existing models. Detailed Implementation
[0059] The technical solution of the present invention will be further explained and described below with reference to the accompanying drawings.
[0060] like Figure 1 This is a schematic diagram of the steel plate surface defect detection system based on spatiotemporal mutual attention and sparse spatiotemporal awareness attention according to the present invention. The steel plate surface defect detection system includes: a steel plate image acquisition module, a shallow feature extraction module, a hybrid deformable convolution module, a multi-scale feature stitching module, a spatiotemporal collaborative feature fusion module, a cross-attention module, a feature stitching module, and an Informer decoder. The steel plate image acquisition module is used to acquire images of the steel plate surface; the shallow feature extraction module is used to extract shallow semantic features of the steel plate surface from the images; the hybrid deformable convolution module is used to perform multi-scale feature extraction on the extracted shallow semantic features of the steel plate surface; and the multi-scale feature stitching module is used to stitch together multiple features. The invention employs a multi-scale feature fusion module to extract spatial-temporal collaborative features from the spliced multi-scale features using a splice-temporal mutual attention mechanism, and then fuses these features with the spliced multi-scale features. A cross-attention module extracts the intersection of high-resolution texture features and low-resolution semantic features from the multi-scale features using a cross-attention mechanism. A feature splicing module splices the fused and intersecting features. The Informer decoder extracts sparsely focused features from the spliced features using a sparse splice-temporal-aware self-attention mechanism, and predicts the type of surface defects based on these features. This invention enhances spatial feature extraction capabilities through multi-scale deformable convolution, combines this with the splice-temporal collaborative feature fusion module to extract spatial-temporal collaborative features, and introduces a sparse splice-temporal-aware self-attention mechanism in the Informer decoder to focus on high-confidence regions on the steel plate surface, thereby improving the accuracy and efficiency of surface defect detection.
[0061] The shallow feature extraction module in this invention includes a block embedding unit, a layer normalization unit, and a layer normalization unit connected in sequence. The convolutional unit and the block embedding unit are used to divide the acquired steel plate surface image into non-overlapping image patches, and embed all image patches into a fixed-dimensional feature vector through linear mapping; in order to reduce the difference in feature distribution, the layer normalization unit is used to normalize the feature vector. Convolutional units are used to extract shallow semantic features of the steel plate surface from normalized feature vectors.
[0062] The hybrid deformable convolution module in this invention includes: Deformable convolutional unit Deformable convolutional units and Deformable convolutional units have kernels that adaptively shift position based on defect morphology to flexibly model multi-scale defects. Deformable convolutional units utilize layer scaling to scale the extracted shallow semantic features of the steel plate surface channel by channel to stabilize the feature distribution and avoid gradient vanishing or exploding problems. Then, they utilize... Deformable convolutional layers extract local detail enhancement features from the surface of steel plates; Deformable convolutional units utilize adaptive activation functions to perform nonlinear transformations on local detail enhancement features, thereby improving the ability of steel plate surface defect detection systems to characterize complex defect patterns. Then, they utilize... Deformable convolutional layers extract mesoscale features from the surface of steel plates; Deformable convolutional units utilize layer normalization to normalize mesoscale features, mitigating training instability caused by uneven distribution. Then, they utilize... Deformable convolutional layers extract global features from the steel plate surface under a larger receptive field. Through this hybrid convolutional structure that expands stepwise from "local to mesoscale to global," and combined with the synergistic effect of layer scaling, adaptive activation functions, and layer normalization, it is possible to simultaneously take into account both subtle defect features and large-scale structural defects, achieving efficient representation and robust modeling of complex defect patterns.
[0063] The multi-scale feature stitching module in this invention is used to stitch together multi-scale features. ,in, These represent local detail enhancement features, mesoscale features, and global features, respectively. This indicates that feature splicing operations are performed along the channel dimension to achieve joint expression of multi-scale information, thereby enhancing the overall feature discrimination ability while preserving the diversity of features in each branch.
[0064] To correct blurred edges and illumination interference, this invention introduces a cross-attention module to extract the intersection of high-resolution texture features and low-resolution semantic features on the steel plate surface using a cross-attention mechanism. The specific process is as follows:
[0065] i: Upsampling is used to obtain high-resolution texture features of the steel plate surface by enhancing local details. It can preserve rich edge and detail information; it obtains low-resolution semantic features of the steel plate surface by downsampling the global features of the steel plate surface. It possesses stronger semantic abstraction capabilities and global context representation;
[0066] ii: High-resolution texture features on the steel plate surface Mapping to query matrix Bond matrix Low-resolution semantic features of steel plate surface Mapping to value matrix The cross-attention mechanism is used to extract the features at the intersection of high-resolution texture features and low-resolution semantic features on the surface of steel plates. :
[0067]
[0068] in, These are the first learnable weight matrix, the second learnable weight matrix, and the third learnable weight matrix, respectively. is the scaling factor for the key matrix. This indicates the transpose operation.
[0069] Through this cross-attention mechanism, the detailed features of the high-resolution branch can be enhanced under the guidance of global semantic information, improving the feature modeling ability under edge blurring and illumination interference, and enhancing the robustness of the characterization and detection of defect areas on the steel plate surface.
[0070] Existing attention fusion structures typically model feature dependencies independently only in the spatial or temporal dimensions, lacking a bidirectional spatial-temporal information interaction mechanism. This results in insufficient ability to characterize temporal variations when dealing with dynamic defects on steel plate surfaces, such as crack propagation and oxidation diffusion. To address this, the spatiotemporal collaborative feature fusion module of this invention utilizes a spatiotemporal mutual attention mechanism to extract spatial-temporal collaborative features, establishing a bidirectional interaction channel between the spatial features extracted by convolution and the embedded temporal series features. The process is as follows:
[0071] i: Perform temporal modeling on the spliced multi-scale features to extract the temporal context features of the steel plate surface. It can capture the evolution pattern of steel plate surface defects over time and extract the temporal context features of the steel plate surface. Mapping to the temporal feature key matrix With time series eigenvalue matrix ;
[0072] ii: Spatial modeling is performed on the spliced multi-scale features using a 1×1 convolution operation to extract the spatial features of the steel plate surface. Enhance spatial texture and boundary features, and integrate them with the temporal context features of the steel plate surface. Alignment within the same feature space, aligning the spatial features of the steel plate surface Mapping to spatial feature query matrix ;
[0073] iii: Transform the temporal feature key matrix Temporal eigenvalue matrix Spatial feature query matrix The spatial-temporal synergistic features of the steel plate surface are integrated:
[0074]
[0075] in, This is a scaling factor for the temporal feature key matrix, used to stabilize the gradient; It is a normalized exponential function.
[0076] The process of fusing the spatial-temporal synergistic features of the steel plate surface with the multi-scale features of the splicing is as follows:
[0077]
[0078] in, Indicates the characteristics of fusion, This represents the multi-scale features of the splicing. express The learnable balance coefficient, express The learnable balance coefficient.
[0079] The spatiotemporal collaborative feature fusion module of the present invention realizes bidirectional feature interaction between "space → time" and "time → space" through mutual attention mechanism, thereby enabling more complete capture of the dynamic changes of steel plate surface defects in the process of morphological evolution and effectively improving the modeling ability of dynamic defects such as crack propagation and oxidation diffusion.
[0080] The feature splicing module in this invention will fuse the features. and cross features By splicing the data, we can obtain the comprehensive features. It can simultaneously preserve the internal dependencies of features and complementary information across resolutions.
[0081] The Informer decoder in this invention includes: a sparse spatiotemporal awareness self-attention mechanism module, a first-layer normalization unit, a feedforward neural network, and a second-layer normalization unit. The sparse spatiotemporal awareness self-attention mechanism module is used to extract sparse focusing features of the steel plate surface from the features spliced by the feature splicing module, and perform spatiotemporal alignment and high-confidence region focusing. The first-layer normalization unit is used to normalize the sparse focusing features of the steel plate surface. The feedforward neural network is used to perform nonlinear transformation on the normalized sparse focusing features. The second-layer normalization unit is used to normalize the nonlinearly transformed features and predict the probability of the defect type on the steel plate surface.
[0082] The process for extracting the sparse focusing features on the surface of the steel plate in this invention is as follows:
[0083] i: Unify the spatial size and temporal length dimensions of the features concatenated by the feature concatenation module, and flatten them into a spatiotemporal token sequence. Map the spacetime token sequence to a spacetime token sequence value matrix. ,in, Denotes the first linear mapping matrix;
[0084] ii: For each spatiotemporal token in the spatiotemporal token sequence, the coordinate information of the query position corresponding to the spatiotemporal token and the time step to which it belongs are included, and the spatiotemporal token is mapped to a spatiotemporal token query matrix and a spatiotemporal token key matrix;
[0085] iii: To reduce complexity and suppress irrelevant regions, a spatiotemporal sparse mask threshold is introduced. For any two query positions in the spatiotemporal token sequence, the sparse attention weight of the query position is calculated, and the sparse attention weight only exists in its local spatial neighborhood and local temporal neighborhood.
[0086] The calculation process for the sparse attention weight of the query position in this invention is as follows:
[0087]
[0088] in, Indicates the query position in the spacetime token sequence and query location Sparse attention weights, Indicates the query position The spatiotemporal token query matrix , Indicates the query position The spacetime token Indicates the corresponding query position The second linear mapping matrix, Indicates the query position The spacetime token key matrix, , Indicates the query position Spacetime token, Indicates the corresponding query position The third linear mapping matrix, This indicates the transpose operation. Indicates the query position The scaling factor of the spacetime token key matrix. Indicates the query position Coordinate information, Indicates the query position Coordinate information, Indicates the query position The time step to which it belongs Indicates the query position The time step to which it belongs This represents the spatial sparse mask threshold in the spatiotemporal sparse mask thresholding. This represents the temporal sparse mask threshold in the spatiotemporal sparse mask threshold;
[0089] iv: Introduce sparse attention weights corresponding to the query position into the spatiotemporal token sequence value matrix to obtain sparse focusing features. ,in, .
[0090] Through the above structural design, the sparse spatiotemporal perception self-attention mechanism further applies local and... Sparsity constraints allow attention to be focused only on key defect areas and key time periods; simultaneously, in terms of comprehensive features... Sparse filtering is performed to effectively suppress noise and illumination interference. This sparse spatiotemporal awareness self-attention mechanism reduces the computational complexity of attention from... Down to It significantly improves computational efficiency while maintaining detection accuracy.
[0091] In one technical solution of the present invention, the steel plate surface images collected in history are used as input, and the types of steel plate surface defects on the collected steel plate surface images are used as labels to train the steel plate surface defect detection system until the dynamic weighted total loss function that integrates normal region loss, defect region loss and neighborhood adaptive loss converges, thus completing the training of the steel plate surface defect detection system.
[0092] To address the challenges of sample sparsity, class imbalance during training, and difficulty in identifying defects in steel plate surface defect detection, this invention proposes a dynamically weighted total loss function. This function adaptively optimizes the learning weights for different regions during training, focusing particularly on defect regions, especially sparse and difficult-to-identify defects, thereby improving the detection accuracy and robustness of the steel plate surface defect detection system. In this invention, the dynamically weighted loss function consists of a normal region loss, a defect region loss, and a neighborhood adaptive loss. To ensure stable learning of the normal region structure during training, the normal region loss dominates in the early stages of training. As training progresses, the weight of the defect region loss gradually increases, especially for sparser and harder-to-identify defect types, which are optimized in particular. Furthermore, based on the severity of different defect types, weights are applied to each type of defect to further improve the steel plate surface defect detection system's ability to identify more difficult defect types. Moreover, to enhance the generalization ability of the steel plate surface defect detection system under different production environments, lighting conditions, or steel grades, a neighborhood adaptive loss is introduced to align the feature distribution between the original training data and the Gaussian blurred training data, thereby reducing the performance degradation caused by neighborhood offset.
[0093] The dynamic weighted total loss function in this invention Represented as:
[0094]
[0095] in, The loss function represents the normal area of the steel plate surface. , This represents the number of samples from the normal area of the steel plate surface. for index, For the first Predicted values for samples from normal areas on the surface of a steel plate. For the first Labels on samples from normal areas of a steel plate surface; The loss function for the defect region on the steel plate surface. , This represents the total number of surface defect types on the steel plate. express index, For the first Number of samples of surface defects in steel plates for index, For the first The severity of surface defects in steel plates. For the first The first sample of the surface defect area of the steel plate Predicted values of surface defects in steel plates For the first The sample from the surface defect area of the steel plate belongs to the first... Labeling of surface defects on steel plates for The adjustment coefficient; This represents the domain-adaptive loss function. , This represents the mean of all predicted values for samples from normal areas and defective areas on the steel plate surface. This represents the covariance of all predicted values for samples from normal areas and samples from defective areas on the steel plate surface. This represents the mean of all predicted values of a pseudo-sample constructed using Gaussian blurring with the domain enhancement operator, combining samples from normal areas and defective areas of the steel plate surface. This represents the covariance of all predicted values of a pseudo-sample constructed using Gaussian blurring with the domain enhancement operator, by combining samples from normal areas and defective areas of the steel plate surface. Denotes the Frobenius norm. for The adjustment coefficient.
[0096] This invention, through the design of the aforementioned dynamic weighted loss function, enables the steel plate surface defect detection system to stably learn the features of normal regions in the early stages of training, while automatically increasing the learning weight of sparse and severely defective regions in the later stages of training, ensuring that these difficult-to-identify defects are fully optimized. At the same time, by combining the aforementioned multi-scale convolution and spatiotemporal collaborative attention to extract key features, it effectively captures the detailed information of steel plate surface defects, thereby achieving high-precision and robust defect detection and adapting to the actual application needs of different production environments.
[0097] The steel plate surface defect detection system based on spatiotemporal mutual attention and sparse spatiotemporal awareness attention of this invention was validated using the NEU-DET industrial defect dataset. This dataset includes six common steel surface defects: patches, oxidation, inclusions, cracks, pitting, and scratches. In the experiment, this invention was tested against mainstream detection models such as Faster R-CNN, SSD, YOLOv8n, YOLOv10n, and YOLOv11n on the same dataset. Specific comparison results are shown below. Figure 2 As shown in the figure, the results show that the average accuracy mAP of the present invention reaches 83.6%, which is higher than the detection performance of the comparison model, especially in the detection of complex defects such as scratches and inclusions. At the same time, the present invention maintains high accuracy while taking into account detection speed, which can meet the needs of real-time detection in industrial sites.
[0098] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should be considered within the scope of protection of the present invention.
Claims
1. A steel plate surface defect detection system based on spatiotemporal mutual attention and sparse spatiotemporal perception attention, characterized in that, include: The system includes a steel plate image acquisition module, a shallow feature extraction module, a hybrid deformable convolution module, a multi-scale feature stitching module, a spatiotemporal collaborative feature fusion module, a cross-attention module, a feature stitching module, and an Informer decoder. The steel plate image acquisition module is used to acquire images of the steel plate surface. The shallow feature extraction module is used to extract shallow semantic features of the steel plate surface from the steel plate surface image; The hybrid deformable convolution module is used to perform multi-scale feature extraction on the extracted shallow semantic features of the steel plate surface; The multi-scale feature stitching module is used to stitch together multi-scale features; The spatiotemporal collaborative feature fusion module is used to extract the spatial-temporal collaborative features of the steel plate surface from the spliced multi-scale features using a spatiotemporal mutual attention mechanism, and to fuse the spatial-temporal collaborative features of the steel plate surface with the spliced multi-scale features; The cross-attention module is used to extract the features at the intersection of high-resolution texture features and low-resolution semantic features on the surface of the steel plate using a cross-attention mechanism; The feature splicing module is used to splice fused features and intersecting features; The Informer decoder is used to extract sparse focused features of the steel plate surface from the spliced features using a sparse spatiotemporal awareness self-attention mechanism, and to predict the type of defects on the steel plate surface based on the sparse focused features of the steel plate surface. The process by which the spatiotemporal collaborative feature fusion module extracts the spatial-temporal collaborative features of the steel plate surface using the spatiotemporal mutual attention mechanism is as follows: i: Perform temporal modeling on the spliced multi-scale features to extract the temporal context features of the steel plate surface. and the time context features of the steel plate surface Mapping to the temporal feature key matrix With time series eigenvalue matrix ; ii: Spatial modeling of the spliced multi-scale features to extract spatial features from the steel plate surface. and the spatial features of the steel plate surface Mapping to spatial feature query matrix ; iii: Transform the temporal feature key matrix Temporal eigenvalue matrix Spatial feature query matrix The spatial-temporal synergistic features of the steel plate surface are integrated: in, The scaling factor for the time series feature key matrix. For normalized exponential functions, Indicates the transpose operation; The process of extracting sparse focusing features from the surface of a steel plate is as follows: i: Unify the dimensions of the features spliced by the feature splicing module, flatten them into a spatiotemporal token sequence, and map the spatiotemporal token sequence into a spatiotemporal token sequence value matrix; ii: For each spatiotemporal token in the spatiotemporal token sequence, the coordinate information of the query position corresponding to the spatiotemporal token and the time step to which it belongs are included, and the spatiotemporal token is mapped to a spatiotemporal token query matrix and a spatiotemporal token key matrix; iii: For any two query positions in the spatiotemporal token sequence, introduce a spatiotemporal sparse mask threshold and calculate the sparse attention weights of the query positions: in, Indicates the query position in the spacetime token sequence and query location Sparse attention weights, Indicates the query position The spatiotemporal token query matrix Indicates the query position The spacetime token key matrix, This indicates the transpose operation. Indicates the query position The scaling factor of the spacetime token key matrix. Indicates the query position Coordinate information, Indicates the query position Coordinate information, Indicates the query position The time step to which it belongs Indicates the query position The time step to which it belongs This represents the spatial sparse mask threshold in the spatiotemporal sparse mask thresholding. This represents the temporal sparse mask threshold in the spatiotemporal sparse mask threshold; iv: Introduce sparse attention weights corresponding to the query position into the spatiotemporal token sequence value matrix to obtain sparse focusing features.
2. The steel plate surface defect detection system based on spatiotemporal mutual attention and sparse spatiotemporal perception attention according to claim 1, characterized in that, The shallow feature extraction module includes, in sequence, a block embedding unit, a layer normalization unit, and... Convolutional unit; The block embedding unit is used to divide the acquired steel plate surface image into non-overlapping image patches, and embed all image patches into a fixed-dimensional feature vector through linear mapping; The layer normalization unit is used to normalize the feature vector; The Convolutional units are used to extract shallow semantic features of the steel plate surface from normalized feature vectors.
3. The steel plate surface defect detection system based on spatiotemporal mutual attention and sparse spatiotemporal perception attention according to claim 1, characterized in that, The hybrid deformable convolutional module includes: Deformable convolutional unit Deformable convolutional units and Deformable convolutional unit; The Deformable convolutional units utilize layer scaling to perform channel-by-channel scaling on the extracted shallow semantic features of the steel plate surface, and then utilize... Deformable convolutional layers extract local detail enhancement features from the surface of steel plates; The Deformable convolutional units utilize adaptive activation functions to perform nonlinear transformations on local detail enhancement features, and then utilize... Deformable convolutional layers extract mesoscale features from the surface of steel plates; The Deformable convolutional units normalize mesoscale features using layer normalization, and then utilize... Deformable convolutional layers extract global features from the surface of steel plates.
4. The steel plate surface defect detection system based on spatiotemporal mutual attention and sparse spatiotemporal perception attention according to claim 3, characterized in that, The cross-attention module is used to extract the features at the intersection of high-resolution texture features and low-resolution semantic features on the surface of the steel plate using a cross-attention mechanism. The process is as follows: i: Upsampling is used to obtain high-resolution texture features of the steel plate surface by enhancing local details. The global features of the steel plate surface are downsampled to obtain low-resolution semantic features of the steel plate surface. ; ii: High-resolution texture features on the steel plate surface Mapping to query matrix Bond matrix Low-resolution semantic features of steel plate surface Mapping to value matrix The cross-attention mechanism is used to extract the features at the intersection of high-resolution texture features and low-resolution semantic features on the surface of steel plates. : in, These are the first learnable weight matrix, the second learnable weight matrix, and the third learnable weight matrix, respectively. is the scaling factor for the key matrix. This indicates the transpose operation.
5. The steel plate surface defect detection system based on spatiotemporal mutual attention and sparse spatiotemporal perception attention according to claim 1, characterized in that, The process of fusing the spatial-temporal synergistic features of the steel plate surface with the multi-scale features of the splicing is as follows: in, Indicates the characteristics of fusion, This represents the multi-scale features of the splicing. express The learnable balance coefficient, express The learnable balance coefficient.
6. The steel plate surface defect detection system based on spatiotemporal mutual attention and sparse spatiotemporal perception attention according to claim 1, characterized in that, The Informer decoder includes: a sparse spatiotemporal awareness self-attention mechanism module, a first-layer normalization unit, a feedforward neural network, and a second-layer normalization unit; The sparse spatiotemporal perception self-attention mechanism module is used to extract sparse focused features on the surface of the steel plate from the features spliced by the feature splicing module; The first normalization unit is used to normalize the sparse focusing features on the surface of the steel plate; The feedforward neural network is used to perform a nonlinear transformation on the normalized sparse focused features; The second normalization unit is used to normalize the features of the nonlinear transformation and predict the probability of the type of defect on the steel plate surface.
7. The steel plate surface defect detection system based on spatiotemporal mutual attention and sparse spatiotemporal perception attention according to claim 1, characterized in that, Using historically collected images of steel plate surfaces as input, and the types of steel plate surface defects on these images as labels, the steel plate surface defect detection system is trained until the total loss function, which integrates normal region loss, defect region loss, and neighborhood adaptive loss, converges, thus completing the training of the steel plate surface defect detection system.
8. A steel plate surface defect detection system based on spatiotemporal mutual attention and sparse spatiotemporal perception attention according to claim 7, characterized in that, The total loss function Represented as: in, The loss function represents the normal area of the steel plate surface. , This represents the number of samples from the normal area of the steel plate surface. for index, For the first Predicted values for samples from normal areas on the surface of a steel plate. For the first Labels on samples from normal areas of a steel plate surface; The loss function for the defect region on the steel plate surface. , This represents the total number of surface defect types on the steel plate. express index, For the first Number of samples of surface defects in steel plates for index, For the first The severity of surface defects in steel plates. For the first The first sample of the surface defect area of the steel plate Predicted values of surface defects in steel plates For the first The sample from the surface defect area of the steel plate belongs to the first... Labeling of surface defects on steel plates for The adjustment coefficient; This represents the domain-adaptive loss function. , This represents the mean of all predicted values for samples from normal areas and defective areas on the steel plate surface. This represents the covariance of all predicted values for samples from normal areas and samples from defective areas on the steel plate surface. This represents the mean of all predicted values of a pseudo-sample constructed using Gaussian blurring with the domain enhancement operator, combining samples from normal areas and defective areas of the steel plate surface. This represents the covariance of all predicted values of a pseudo-sample constructed using Gaussian blurring with the domain enhancement operator, by combining samples from normal areas and defective areas of the steel plate surface. Denotes the Frobenius norm. for The adjustment coefficient.
Citation Information
Patent Citations
Strip steel surface defect detection method based on attention mechanism and improved UNet network
CN119722666A
Steel surface microscopic crack feature extraction method, device and equipment and storage medium
CN120510135A