A steel plate surface defect detection method, device and electronic equipment

By employing multi-device acquisition and feature fusion technology, the problems of image distortion and multi-scale adaptability in steel plate surface defect detection have been solved, achieving high-precision defect detection that is suitable for various application scenarios.

CN122289124APending Publication Date: 2026-06-26ZHEJIANG WANLI UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG WANLI UNIV
Filing Date
2026-02-11
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

In existing methods for detecting defects on steel plate surfaces, single acquisition devices suffer from image detail distortion when illumination is unstable, and images from multiple acquisition devices are simply stitched together and fused, which is insufficient to adapt to defects at multiple scales, resulting in decreased detection accuracy and difficulty in adapting to various application scenarios.

Method used

Industrial visible light cameras, infrared imaging equipment, lidar equipment, and spectral sensors are used to simultaneously collect data on the surface of steel plates. Candidate defect regions are determined by feature tensor normalization and mapping relationships. Cross-modal feature alignment and weighted fusion are performed using fusion loss functions and joint scoring functions to generate high-precision defect masks.

Benefits of technology

It achieves precise alignment of cross-modal features, improves the accuracy and sensitivity of defect detection, adapts to various application scenarios, effectively avoids noise mode interference, and meets the needs of multi-scale defect detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122289124A_ABST
    Figure CN122289124A_ABST
Patent Text Reader

Abstract

This disclosure provides a method, apparatus, and electronic device for detecting surface defects in steel plates. The method involves simultaneously acquiring multiple surface data points of the steel plate under test using different acquisition devices. For any given surface data point, a first modal feature tensor is extracted and normalized to obtain a second modal feature tensor. Based on a preset standard modal feature tensor and a first mapping relationship, first candidate defect regions are determined for all second modal feature tensors. Second candidate defect regions are selected from all first candidate defect regions. Third candidate defect regions are obtained based on the cosine similarity and fusion loss function of the second candidate defect regions. All third candidate defect regions are weighted and fused to obtain a fused defect region. The target defect is determined based on the candidate defect mask and a joint scoring function. This improves the accuracy of steel surface defect detection and can adapt to the detection needs of various complex industrial scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the fields of computer vision and industrial nondestructive testing technology, and in particular to a method, apparatus and electronic device for detecting defects on steel surfaces. Background Technology

[0002] Currently, defect detection on steel plate surfaces mostly employs either single-device image acquisition or simple image stitching and fusion from multiple devices. However, images acquired by a single device are prone to detail distortion under unstable lighting conditions, leading to decreased detection accuracy. Multiple-device images undergo only simple global stitching or weighted fusion, lacking regional-level cross-modal semantic alignment constraints, making them susceptible to noise mode interference and difficult to adapt to multi-scale defects. Local failures in some modes directly lower overall detection accuracy. Thus, the accuracy of steel plate surface defect identification decreases, making it difficult to adapt to various application scenarios. Summary of the Invention

[0003] This disclosure provides a method, apparatus, and electronic device for detecting surface defects in steel plates, which to some extent solves the problem of declining accuracy in existing steel plate surface defect identification and difficulty in adapting to various application scenarios.

[0004] According to one aspect of this disclosure, a method for detecting surface defects in steel plates is provided. The method includes: simultaneously acquiring multiple surface data of the steel plate to be tested using different acquisition devices; the different acquisition devices include: an industrial visible light camera, an infrared imaging device, a lidar device, and a spectral sensor; for any given surface data, extracting a first modal feature tensor of the surface data and performing normalization processing to obtain a second modal feature tensor; the second modal feature tensor is used to indicate the first modal feature tensor with uniform dimension and scale; based on a preset standard modal feature tensor and a first mapping relationship, determining a first candidate defect region for all second modal feature tensors; the standard modal feature tensor is any one of the second modal feature tensors; the first mapping relationship includes: a second mapping relationship between the second modal feature tensor and the steel plate to be tested. The system defines a third mapping relationship between multiple second-modal feature tensors; it selects second-candidate defect regions from all first-candidate defect regions that satisfy a preset first intersection-union ratio (IU), and determines the cosine similarity of each second-candidate defect region; based on cosine similarity and a fusion loss function, it aligns all second-candidate defect regions to obtain third-candidate defect regions; the fusion loss function is determined based on detection loss, masking loss, and cross-modal consistency loss; it performs weighted fusion on all third-candidate defect regions to obtain fused defect regions; it inputs the fused defect regions and second-candidate defect regions into a segmentation model to generate multiple candidate defect masks; it determines the target defect based on the candidate defect masks and a joint scoring function; the joint scoring function is determined based on the second IU and the semantic similarity between the candidate defect masks and the fused defect regions.

[0005] Furthermore, according to one aspect of the method of this disclosure, the target defect includes at least one of the following: scratches, indentations, bubbles, fine cracks, pinholes, and folds.

[0006] Furthermore, according to one aspect of the method of this disclosure, the first candidate defect region of all second modal feature tensors is determined based on a preset standard modal feature tensor and a first mapping relationship, including: selecting any second modal feature tensor as a standard modal feature tensor; determining candidate defect boxes based on a feature pyramid network; mapping the candidate defect boxes to all second modal feature tensors other than the standard modal feature tensor based on the candidate defect boxes and a third mapping relationship; and associating the mapped second modal feature tensors with the steel plate to be tested based on the second mapping relationship to obtain the first candidate defect region.

[0007] Furthermore, according to one aspect of the method of this disclosure, second candidate defect regions that satisfy a preset first crossover-union ratio are selected from all first candidate defect regions, and the cosine similarity of each second candidate defect region is determined, including: determining the crossover-union ratio of each first candidate defect region with a standard defect mask; determining all first candidate defect regions whose crossover-union ratio is greater than or equal to the first crossover-union ratio as second candidate defect regions; and obtaining the cosine similarity of all any two second candidate defect regions based on a cosine similarity algorithm.

[0008] Furthermore, according to one aspect of the method disclosed herein, a third candidate defect region is obtained by aligning all second candidate defect regions based on cosine similarity and a fusion loss function, including: for any two second candidate defect regions, determining the feature consistency error between the second candidate defect regions based on cosine similarity to obtain a cross-modal consistency loss; obtaining a detection loss; using the detection loss to constrain the localization accuracy and category recognition accuracy of the second candidate defect regions; obtaining a masking loss; using the masking loss to constrain the contour segmentation accuracy of the second candidate defect regions; weightedly fusing the cross-modal consistency loss, detection loss, and masking loss to obtain a fusion loss function; and obtaining the third candidate defect region when the fusion loss value of all second candidate defect regions is minimized.

[0009] Furthermore, according to one aspect of the method disclosed herein, the segmentation model includes: a convolutional neural network model, a converter segmentation network model, and a general segmentation model.

[0010] Furthermore, according to one aspect of the method of this disclosure, the fused defect region and the second candidate defect region are input into a segmentation model to generate multiple candidate defect masks, including: determining a generated bounding box cue, a point cue, and an edge cue based on the second candidate region; and inputting the generated bounding box cue, the point cue, the edge cue, and the fused defect model into the segmentation model to obtain the candidate defect masks.

[0011] Furthermore, according to one aspect of the method disclosed herein, a target defect is determined based on a candidate defect mask and a joint scoring function, including: determining a second intersection-union ratio (IUU) between the candidate defect mask and a standard defect mask; obtaining semantic similarity; weightedly fusing the second IUU and the semantic similarity to obtain a joint scoring function; inputting each candidate defect mask into the joint scoring function to obtain corresponding score values; and determining the corresponding candidate defect mask as the target defect when the score value is the maximum.

[0012] According to another aspect of this disclosure, a steel plate surface defect detection device is provided. The device includes: an acquisition unit, used to simultaneously acquire multiple surface data of the steel plate to be tested using different acquisition devices; the different acquisition devices include: an industrial visible light camera, an infrared imaging device, a lidar device, and a spectral sensor; an extraction unit, used to extract a first modal feature tensor of any surface data and perform normalization processing to obtain a second modal feature tensor; the second modal feature tensor is used to indicate the first modal feature tensor with uniform dimension and scale; a first determination unit, used to determine the first candidate defect region of all second modal feature tensors based on a preset standard modal feature tensor and a first mapping relationship; the standard modal feature tensor is any one of the second modal feature tensors; the first mapping relationship includes: a second mapping relationship between the second modal feature tensor and the steel plate to be tested, and multiple second modal feature tensors. The third mapping relationship between the two modal feature tensors; the second determination unit, used to filter second candidate defect regions that meet the preset first intersection-union ratio from all first candidate defect regions, and determine the cosine similarity of each second candidate defect region; the alignment unit, used to align all second candidate defect regions to obtain third candidate defect regions based on cosine similarity and fusion loss function; the fusion loss function is determined based on detection loss, masking loss and cross-modal consistency loss; the fusion unit, used to perform weighted fusion on all third candidate defect regions to obtain fused defect regions; the generation unit, used to input the fused defect regions and second candidate defect regions into the segmentation model to generate multiple candidate defect masks; the third determination unit, used to determine the target defect based on the candidate defect masks and the joint scoring function; the joint scoring function is determined based on the second intersection-union ratio, and the semantic similarity between the candidate defect masks and the fused defect regions.

[0013] According to another aspect of this disclosure, an electronic device is provided, comprising: a memory for storing computer-readable instructions; and a processor for executing the computer-readable instructions, causing the electronic device to perform the method as described in any embodiment of one aspect.

[0014] This disclosure provides a method, apparatus, and electronic device for detecting surface defects in steel plates. The method involves simultaneously acquiring multiple surface data points of the steel plate under test using different acquisition devices, including industrial visible light cameras, infrared imaging devices, lidar devices, and spectral sensors. For any given surface data point, a first modal feature tensor is extracted and normalized to obtain a second modal feature tensor. This second modal feature tensor indicates the first modal feature tensor with uniform dimension and scale. Based on a preset standard modal feature tensor and a first mapping relationship, first candidate defect regions for all second modal feature tensors are determined. The standard modal feature tensor is any one of the second modal feature tensors. The first mapping relationship includes a second mapping relationship between the second modal feature tensors and the steel plate under test, and multiple second modal feature tensors... The third mapping relationship between quantities; second candidate defect regions that satisfy the preset first intersection-union ratio are selected from all first candidate defect regions, and the cosine similarity of each second candidate defect region is determined; based on the cosine similarity and the fusion loss function, all second candidate defect regions are aligned to obtain third candidate defect regions; the fusion loss function is determined based on detection loss, masking loss and cross-modal consistency loss; all third candidate defect regions are weighted and fused to obtain fused defect regions; the fused defect regions and second candidate defect regions are input into the segmentation model to generate multiple candidate defect masks; the target defect is determined based on the candidate defect masks and the joint scoring function; the joint scoring function is determined based on the second intersection-union ratio and the semantic similarity between the candidate defect masks and the fused defect regions. In contrast to existing single-modal detection methods, which suffer from weak anti-interference capabilities, multi-modal detection methods that only involve simple global fusion, and a lack of region-level alignment and adaptive anti-interference mechanisms, this disclosure utilizes multi-device acquisition and achieves precise alignment of cross-modal features at the defect region granularity. It simultaneously constrains detection accuracy, segmentation accuracy, and cross-modal consistency through a fusion loss function. Furthermore, by employing a dual-intersection comparison and joint scoring mechanism, it effectively avoids noise mode interference, adapts to multi-scale defect detection needs, and improves defect identification sensitivity. In summary, the technical solution provided by this disclosure can improve the accuracy of steel surface defect detection and can be adapted to various application scenarios.

[0015] It should be understood that both the foregoing general description and the following detailed description are exemplary and intended to provide further illustration of the claimed technology. Attached Figure Description

[0016] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0017] Figure 1 A schematic flowchart of a method for detecting surface defects in steel plates provided in this embodiment of the present disclosure; Figure 2 A schematic flowchart illustrating a complete defect detection method provided in this embodiment of the disclosure; Figure 3 This is a structural block diagram of a steel plate surface defect detection device provided in an embodiment of the present disclosure; Figure 4 This is a hardware block diagram of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this disclosure more apparent, exemplary embodiments according to this disclosure will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this disclosure, and not all embodiments of this disclosure. It should be understood that this disclosure is not limited to the exemplary embodiments described herein.

[0019] Currently, defect detection on steel plate surfaces mostly employs either single-device image acquisition or simple image stitching and fusion from multiple devices. However, images acquired by a single device are prone to detail distortion under unstable lighting conditions, leading to decreased detection accuracy. Multiple-device images undergo only simple global stitching or weighted fusion, lacking regional-level cross-modal semantic alignment constraints, making them susceptible to noise mode interference and difficult to adapt to multi-scale defects. Local failures in some modes directly lower overall detection accuracy. Thus, the accuracy of steel plate surface defect identification decreases, making it difficult to adapt to various application scenarios.

[0020] To address the aforementioned issues, this disclosure provides a method for detecting surface defects in steel plates. Compared to existing single-modal detection methods, which suffer from weak anti-interference capabilities, limited multi-modal detection methods that only involve simple global fusion, and lack regional alignment and adaptive anti-interference mechanisms, this disclosure utilizes multi-device acquisition and achieves precise alignment of cross-modal features at the defect region granularity. It simultaneously constrains detection accuracy, segmentation accuracy, and cross-modal consistency through a fusion loss function. Furthermore, by employing a dual-intersection comparison and joint scoring mechanism, it effectively avoids noise mode interference, adapts to multi-scale defect detection requirements, and improves defect identification sensitivity. In summary, the technical solution provided by this disclosure can improve the accuracy of steel surface defect detection and is adaptable to various application scenarios.

[0021] This disclosure provides a method for detecting surface defects in steel plates. Please refer to... Figure 1 , Figure 1 This is a schematic flowchart illustrating a method for detecting surface defects in steel plates, provided in an embodiment of this disclosure. Figure 1 As shown, the method includes: In step S101, multiple surface data of the steel plate under test are simultaneously acquired using different acquisition devices; the different acquisition devices include: industrial visible light camera, infrared imaging device, lidar device and spectral sensor; In step S102, for any surface data, the first modal feature tensor of the surface data is extracted and normalized to obtain the second modal feature tensor; the second modal feature tensor is used to indicate the first modal feature tensor with uniform dimension and scale. In step S103, based on the preset standard modal feature tensor and the first mapping relationship, the first candidate defect region of all second modal feature tensors is determined; the standard modal feature tensor is any second modal feature tensor; the first mapping relationship includes: the second mapping relationship between the second modal feature tensor and the steel plate to be tested, and the third mapping relationship between multiple second modal feature tensors; In step S104, second candidate defect regions that satisfy a preset first intersection-union ratio are selected from all first candidate defect regions, and the cosine similarity of each second candidate defect region is determined. In step S105, based on cosine similarity and fusion loss function, all second candidate defect regions are aligned to obtain the third candidate defect region; the fusion loss function is determined based on detection loss, masking loss and cross-modal consistency loss. In step S106, all third candidate defect regions are weighted and fused to obtain a fused defect region; In step S107, the fused defect region and the second candidate defect region are input into the segmentation model to generate multiple candidate defect masks; In step S108, the target defect is determined based on the candidate defect mask and the joint scoring function; the joint scoring function is determined based on the second intersection-union ratio and the semantic similarity between the candidate defect mask and the fused defect region.

[0022] In this disclosure, the acquisition device can be understood as a sensing device used to acquire different modal physical information of the steel plate surface, capable of simultaneously acquiring multi-dimensional data such as visual, thermal, three-dimensional structure, and spectral data of the steel plate surface. Different acquisition devices disclosed in this disclosure include: industrial visible light cameras, infrared imaging devices, lidar devices, and spectral sensors. Specifically, the industrial visible light camera acquires RGB visual images of the steel plate surface, reflecting surface texture, color, and geometric defect characteristics; the infrared imaging device acquires thermal imaging images of the steel plate surface, reflecting the temperature distribution and thermal characteristic defects of the surface and shallow interior; the lidar device acquires point cloud data of the steel plate surface, reflecting the surface three-dimensional morphology and three-dimensional defect characteristics; and the spectral sensor acquires spectral data of the steel plate surface, reflecting the surface material composition and material defect characteristics.

[0023] In this disclosure, the first modal feature tensor can be understood as tensor data composed of multi-scale feature maps directly extracted from the original data of the steel plate surface in a single modality through a feature extraction network. It is a deep feature representation of the original data of the steel plate surface in that modality. The first modal feature tensors of different modalities have different channel dimensions and feature scales.

[0024] In this disclosure, normalization can be understood as the operation of numerically standardizing the first modal feature tensors of each modality. By eliminating the differences in dimensions and numerical ranges between different modal feature tensors, all modal features are placed in the same numerical space. The second modal feature tensor obtained after normalization can be understood as the first modal feature tensor being normalized and then mapped to tensor data with a unified channel dimension and feature scale through 1×1 convolution or linear transformation. It is a standardized representation of each modal feature in a unified feature space, realizing the unification of the dimension and scale of different modal features.

[0025] In this disclosure, the standard modal feature tensor can be understood as any modal feature tensor selected from all second modal feature tensors, which serves as a reference benchmark for generating cross-modal defect candidate regions. The second modal feature tensors of the remaining modes are used as a reference for mapping and aligning candidate regions.

[0026] In this disclosure, the first mapping relationship can be understood as a mapping system used to realize the spatial position association between multimodal feature tensors and between feature tensors and the physical coordinate system of the steel plate. It is the core basis for cross-modal candidate region alignment and defect physical location restoration. The second mapping relationship can be understood as a spatial mapping relationship that associates the image coordinate system of the second-modal feature tensor with the physical coordinate system of the steel plate under test, enabling the conversion of the image coordinates of the defect region in the feature tensor to the actual physical coordinates of the steel plate. The third mapping relationship can be understood as a spatial mapping relationship that associates the image coordinate systems of second-modal feature tensors of different modalities, enabling accurate mapping of the defect candidate region in a certain modality to all other modal feature tensors.

[0027] In this disclosure, the first candidate defect region can be understood as the candidate defect box obtained by the defect candidate region generation network based on the standard modal feature tensor, and then mapped to all second modal feature tensors through the third mapping relationship. After being associated with the physical location of the steel plate through the second mapping relationship, the suspected defect regions in each modal feature tensor are initially determined.

[0028] In this disclosure, the first crossover ratio (CRR) can be understood as the CRR threshold used to screen valid first candidate defect regions, and is the CRR judgment criterion between the first candidate defect region and the manually labeled standard defect mask. The second candidate defect region can be understood as a valid suspected defect region selected from all first candidate defect regions whose CRR with the standard defect mask is greater than or equal to the first CRR.

[0029] In this disclosure, the fusion loss function can be understood as a joint loss function used to constrain the accuracy of cross-modal feature alignment, defect detection, and defect segmentation, and is the core optimization objective of multi-task joint training. Specifically, the detection loss constrains the localization accuracy and category recognition accuracy of the defect region, the masking loss constrains the contour segmentation accuracy of the defect region, and the cross-modal consistency loss constrains the consistency of feature representation of the same physical defect region under different modalities. These three are weighted and fused to form the fusion loss function.

[0030] In this disclosure, the third candidate defect region can be understood as the effective defect region obtained after feature alignment by training all the second candidate defect regions end-to-end through a fusion loss function, so that the feature representations of the same physical defect region tend to be consistent under different modalities.

[0031] In this disclosure, the fused defect region can be understood as a unified defect region representation that integrates the advantages of multimodal features after learning-weighted fusion of all third candidate defect regions, thus combining the feature descriptions of each modality for the defect region.

[0032] In this disclosure, the segmentation model can be understood as a deep learning model used to perform fine contour segmentation of the fused defect region and generate a defect mask. It can output the accurate contour of the defect based on the feature representation and location cues of the defect region. The segmentation model disclosed includes: a convolutional neural network model, a converter segmentation network model, and a general segmentation model. The convolutional neural network model, based on convolution operations, can achieve fine segmentation of the defect contour; the converter segmentation network model, based on a self-attention mechanism, extracts global features and is adapted to the segmentation of complex-shaped defects; the general segmentation model can quickly generate high-quality candidate defect masks based on cues such as points, boxes, and edges.

[0033] In this disclosure, the candidate defect mask can be understood as multiple defect contour candidate results output by the segmentation model after inputting the fused defect region and the second candidate defect region into the segmentation model, which is a pixel-level contour representation of the defect region.

[0034] In this disclosure, the joint scoring function can be understood as a comprehensive scoring function constructed based on the second intersection-union ratio and semantic similarity, used to screen the optimal defect mask, which can evaluate the candidate defect mask from both geometric accuracy and semantic consistency dimensions.

[0035] In this disclosure, the target defect can be understood as the actual defect on the steel plate surface corresponding to the optimal defect mask with the highest score (not lower than a preset threshold) obtained from all candidate defect masks through a joint scoring function. This is the steel plate surface defect ultimately detected and identified by this method. The target defects in this disclosure include at least one of the following: scratches, indentations, bubbles, fine cracks, pinholes, and folds. Among them, scratches are linear defects formed on the steel plate surface due to friction or abrasion; indentations are dent-like defects formed on the steel plate surface due to external force; bubbles are bulge-like defects formed on the steel plate surface due to the leakage of internal gas; fine cracks are micron-sized crack-like defects appearing on the steel plate surface; pinholes are micron-sized pore-like defects appearing on the steel plate surface; and folds are layer-like defects formed during the steel plate production process due to improper rolling.

[0036] Specifically, determining surface defects in steel plates can include the following steps: First, multimodal surface data of the steel plate under test are simultaneously acquired using multiple acquisition devices. Next, a first modal feature tensor is extracted from the raw data of each modality, and normalized to obtain a second modal feature tensor with unified dimensions and scale. Then, based on the standard modal feature tensor and a first mapping relationship, the first candidate defect regions for each modality are obtained. Furthermore, a second candidate defect region is obtained through a first intersection-union-ratio (IUU) screening process, and cosine similarity is calculated. Next, based on cosine similarity and a fusion loss function, a third candidate defect region with aligned features is obtained. Finally, the third candidate defect regions are weighted and fused to obtain a fused defect region. Then, the fused defect region is input into a segmentation model to generate candidate defect masks. Finally, based on a joint scoring function, the optimal result is selected from the candidate defect masks to determine the target defect.

[0037] The following details how to determine the first candidate defect region, including: Choose any one of the second modal feature tensors as the standard modal feature tensor; Candidate defect boxes are determined based on the feature pyramid network. Based on the candidate defect boxes and the third mapping relationship, the candidate defect boxes are mapped to all second modal feature tensors except for the standard modal feature tensor; Based on the second mapping relationship, the mapped second modal feature tensor is correlated with the position of the steel plate to be tested to obtain the first candidate defect region.

[0038] In this disclosure, the feature pyramid network can be understood as a deep learning network for multi-scale feature extraction and fusion, which can generate candidate defect boxes of different sizes using feature maps of different scales, adapting to the detection needs of multi-scale defects on steel plate surfaces.

[0039] In this disclosure, the candidate defect box can be understood as a rectangular box detected by the feature pyramid network on the standard modal feature tensor, used to define the suspected defect region, and containing the image coordinate information of the suspected defect region.

[0040] In this disclosure, position association can be understood as converting the image coordinates of the candidate defect box in the second modality feature tensor into the actual physical coordinates of the steel plate to be tested through the second mapping relationship, thereby realizing a one-to-one correspondence between the defect region in the feature tensor and the actual position of the steel plate.

[0041] Specifically, determining the first candidate defect region may include the following steps: The first step is to randomly select any modality tensor from all second modal feature tensors with uniform dimensions and scales as the standard modal feature tensor, which serves as the benchmark for generating defect candidate regions. The second step is to input the standard modal feature tensor into the defect detection head built on the feature pyramid network. The detection head then detects and outputs multiple candidate defect boxes of different sizes on the multi-scale feature map, covering the suspected defect areas at different scales. The third step is to map all candidate defect boxes from the image coordinate system of the standard modal feature tensor to the image coordinate system of all other second modal feature tensors according to the pre-defined third mapping relationship, so as to obtain the candidate defect boxes corresponding to each modal feature tensor. The fourth step involves converting the image coordinates of the candidate defect boxes in each modal feature tensor into the physical coordinates of the steel plate under test, based on the pre-defined second mapping relationship. This establishes the association between the candidate defect boxes and the actual position of the steel plate, ultimately yielding the corresponding first candidate defect regions in all second modal feature tensors.

[0042] The following will specifically explain how this disclosure determines the second candidate defect region, including: Determine the intersection-union ratio (IUU) of each first candidate defect region with the standard defect mask; All first candidate defect regions whose cross-union ratio is greater than or equal to the first cross-union ratio are identified as second candidate defect regions; Based on the cosine similarity algorithm, the cosine similarity of any two second candidate defect regions is obtained.

[0043] In this disclosure, the intersection-union ratio can be understood as the calculated result of the intersection-union ratio between the first candidate defect region and the manually annotated standard defect mask. It is a quantitative indicator that measures the degree of overlap between the first candidate defect region and the actual defect region.

[0044] In this disclosure, the cosine similarity algorithm can be understood as an algorithm used to calculate the directional similarity between two feature vectors. By calculating the cosine value of the angle between the feature vectors, the similarity of the feature representations of the second candidate defect region under different modes of the same physical defect region is quantified.

[0045] Specifically, determining the second candidate defect region may include the following steps: The first step is to obtain a standard defect mask that is manually labeled on the feature tensor of the standard mode. This mask is a precise contour representation of the actual defects on the steel plate surface. The second step is to calculate the intersection-union ratio of each first candidate defect region with the corresponding standard defect mask to quantify the overlap between the first candidate defect region and the actual defect. The third step is to set a first crossover ratio threshold, filter out the first candidate defect areas whose crossover ratio is greater than or equal to the threshold, and use them as valid second candidate defect areas, while removing invalid suspected areas with too low overlap. The fourth step is to extract regional features from all second candidate defect regions to obtain the feature vectors corresponding to each second candidate defect region. Based on the cosine similarity algorithm, the cosine similarity between the feature vectors of any two second candidate defect regions under different modes of the same physical defect region is calculated, which provides a basis for subsequent cross-modal consistency loss calculation.

[0046] The following will explain in detail how to obtain the third candidate defect region, including: For any two second candidate defect regions, the feature consistency error between the second candidate defect regions is determined based on cosine similarity, and the cross-modal consistency loss is obtained. Obtain the detection loss; the detection loss is used to constrain the localization accuracy and category recognition accuracy of the second candidate defect region. Obtain the mask loss; the mask loss is used to constrain the contour segmentation accuracy of the second candidate defect region; The cross-modal consistency loss, detection loss, and masking loss are weighted and fused to obtain the fusion loss function; The third candidate defect region is obtained when the fusion loss value of all second candidate defect regions is minimized.

[0047] Specifically, determining the third candidate defect region may include the following steps: The first step is to calculate the consistency error between feature vectors based on the cosine similarity of any two second candidate defect regions under different modalities of the same physical defect region, and to construct the cross-modal consistency loss. The larger the error, the higher the loss value, thus constraining the features of different modalities to tend to be consistent. The second step is to extract the location and category information of each second candidate defect region, calculate the detection loss, and constrain the candidate box regression accuracy and defect category classification accuracy of the defect region. The third step is to extract the contour segmentation information of each second candidate defect region, calculate the mask loss, and constrain the pixel-level contour segmentation accuracy of the defect region. The fourth step involves configuring adjustable weighting coefficients for the cross-modal consistency loss, detection loss, and masking loss, and then fusing the three weighted losses to obtain a fusion loss function, which serves as the overall optimization objective for model training. The fifth step involves inputting all second candidate defect regions into the model for end-to-end training, continuously adjusting the model parameters through backpropagation until the loss value of the fusion loss function reaches its minimum. At this point, the feature representations of the second candidate defect regions of the same physical defect region in different modalities tend to be consistent, and the defect region with aligned features is determined as the third candidate defect region.

[0048] The following will explain in detail how to obtain candidate defect masks, including: Based on the second candidate region, determine the generation of bounding box hints, point hints, and edge hints; The generated bounding box hints, point hints, edge hints, and fused defect regions are input into the segmentation model to obtain candidate defect masks.

[0049] In this disclosure, the generated bounding box hint can be understood as a rectangular bounding box hint generated based on the location information of the second candidate defect region, used to indicate the approximate location range of the defect region to the segmentation model. The point hint can be understood as a key pixel hint selected based on the feature information of the second candidate defect region, used to indicate the core location of the defect region to the segmentation model. The edge hint can be understood as an edge pixel hint extracted based on the contour information of the second candidate defect region, used to indicate the contour boundary of the defect region to the segmentation model.

[0050] Specifically, obtaining candidate defect masks may include the following steps: The first step is to generate a rectangular box surrounding the second candidate defect region based on the image coordinate information of the defect region, as a box cue; select several key pixels at the core position of the defect region as point cue; and extract the edge contour pixels of the defect region as edge cue. The second step is to perform learnable weighted fusion on all third candidate defect regions to obtain a fused defect region that incorporates the advantages of multimodal features, and then extract the unified defect representation features of this region. The third step involves inputting location hints such as box hints, point hints, and edge hints, along with the unified defect representation features of the fused defect region, into the segmentation model to provide the segmentation model with location guidance and feature representation of the defects. The fourth step involves the segmentation model performing pixel-level fine contour segmentation on the fused defect region, and outputting multiple defect contour results with different precision based on the defect features and hints, and determining all contour results as candidate defect masks.

[0051] The following will explain in detail how to obtain the target defect, including: Determine the second intersection-union ratio (CUI) between the candidate defect mask and the standard defect mask; Obtain semantic similarity; The second intersection-union ratio and semantic similarity are weighted and fused to obtain a joint scoring function; Each candidate defect mask is input into the joint scoring function to obtain the corresponding score value; When the score is the highest, the corresponding candidate defect mask is determined as the target defect.

[0052] Specifically, identifying the target defect may include the following steps: The first step is to obtain the standard defect mask with manual annotation, calculate the cross-union ratio between each candidate defect mask and the standard defect mask, define the cross-union ratio as the second cross-union ratio, and quantify the geometric overlap between the candidate defect mask and the actual defect contour. The second step is to perform pooling on the features within the coverage area of ​​each candidate defect mask to obtain the mask semantic feature vector. The semantic similarity between this vector and the unified defect representation features of the fused defect region is calculated to quantify the semantic consistency between the candidate defect mask and the fused defect region. The third step involves configuring preset non-negative weighting coefficients α and β for the second intersection-union ratio and semantic similarity, respectively, and then fusing the two weighted scores to construct a joint scoring function. The function expression is as follows: ,in, The overall score of the candidate defect mask. The weighting coefficients for the second intersection-union ratio (IOU); This represents the intersection area / union area of ​​the candidate mask and the standard mask; The weighting coefficients for semantic similarity; The semantic similarity between the candidate defect mask and the fused defect region; The fourth step involves substituting all candidate defect masks into the joint scoring function to calculate the comprehensive score value corresponding to each candidate defect mask, and sorting the candidate defect masks from high to low score values. The fifth step involves setting a scoring threshold, selecting the candidate defect mask with the highest score value after sorting that is not lower than the threshold, converting the image coordinates of the mask into the physical coordinates of the steel plate to be tested based on the second mapping relationship, restoring the actual location, size, and category of the defect, and determining the actual defect on the surface of the steel plate corresponding to the candidate defect mask as the target defect.

[0053] For example, Figure 2 This is a schematic flowchart illustrating a complete defect detection method provided in an embodiment of this disclosure. Figure 2 As can be seen, the entire process includes: first, acquiring and aligning multi-source steel plate surface data such as RGB and infrared through multimodal data acquisition and registration; mapping different modal features to a unified dimension and scale through modal feature extraction and unified encoding; generating candidate defect regions based on standard modal features and completing cross-modal mapping; achieving feature alignment of the same defect region under different modalities through cross-modal feature consistency constraint learning; performing multimodal feature fusion and defect segmentation on the aligned features to generate candidate defect masks; selecting the optimal defect mask from both geometric accuracy and semantic consistency dimensions based on mask scoring and selection; adapting to the detection needs of defects of different sizes by combining multi-scale detection and evaluation; and finally outputting the results, restoring the physical location, size, and category information of the defects.

[0054] This disclosure also provides a device for detecting defects on the surface of steel plates. Figure 3 This is a structural block diagram of a steel plate surface defect detection device provided in an embodiment of the present disclosure, as shown below. Figure 3 As shown, the steel plate surface defect detection device 300 includes: The acquisition unit 301 is used to simultaneously acquire multiple surface data of the steel plate under test using different acquisition devices; the different acquisition devices include: industrial visible light camera, infrared imaging equipment, lidar equipment and spectral sensor; Extraction unit 302 is used to extract the first modal feature tensor of any surface data and perform normalization processing to obtain the second modal feature tensor; the second modal feature tensor is used to indicate the first modal feature tensor with uniform dimension and scale; The first determining unit 303 is used to determine the first candidate defect region of all second modal feature tensors based on a preset standard modal feature tensor and a first mapping relationship; the standard modal feature tensor is any second modal feature tensor; the first mapping relationship includes: a second mapping relationship between the second modal feature tensor and the steel plate to be tested, and a third mapping relationship between multiple second modal feature tensors; The second determining unit 304 is used to filter second candidate defect regions that satisfy a preset first intersection-union ratio from all first candidate defect regions, and to determine the cosine similarity of each second candidate defect region. Alignment unit 305 is used to align all second candidate defect regions to obtain third candidate defect regions based on cosine similarity and fusion loss function; the fusion loss function is determined based on detection loss, masking loss and cross-modal consistency loss. The fusion unit 306 is used to perform weighted fusion of all third candidate defect regions to obtain a fused defect region. The generation unit 307 is used to input the fused defect region and the second candidate defect region into the segmentation model to generate multiple candidate defect masks. The third determining unit 308 is used to determine the target defect based on the candidate defect mask and the joint scoring function; the joint scoring function is determined based on the second intersection-union ratio and the semantic similarity between the candidate defect mask and the fused defect region.

[0055] In one exemplary embodiment, the third determining unit 308 is specifically used to determine that the target defect includes at least one of the following: scratches, indentations, bubbles, fine cracks, pinholes, and folds.

[0056] In one exemplary embodiment, the first determining unit 303 is specifically used to: select any second modal feature tensor as the standard modal feature tensor; determine candidate defect boxes based on the feature pyramid network; map the candidate defect boxes to all second modal feature tensors except the standard modal feature tensor based on the candidate defect boxes and the third mapping relationship; and associate the mapped second modal feature tensors with the steel plate to be tested based on the second mapping relationship to obtain the first candidate defect region.

[0057] In one exemplary embodiment, the second determining unit 304 is specifically used to: determine the intersection-union ratio (IUR) of each first candidate defect region with the standard defect mask; determine all first candidate defect regions whose IUR is greater than or equal to the first IUR as second candidate defect regions; and obtain the cosine similarity of all any two second candidate defect regions based on the cosine similarity algorithm.

[0058] In one exemplary embodiment, the alignment unit 305 is specifically configured to: for any two second candidate defect regions, determine the feature consistency error between the second candidate defect regions based on cosine similarity to obtain a cross-modal consistency loss; obtain a detection loss; the detection loss is used to constrain the localization accuracy and category recognition accuracy of the second candidate defect regions; obtain a mask loss; the mask loss is used to constrain the contour segmentation accuracy of the second candidate defect regions; perform weighted fusion of the cross-modal consistency loss, detection loss, and mask loss to obtain a fusion loss function; and obtain a third candidate defect region when the fusion loss value of all second candidate defect regions is minimized.

[0059] In one exemplary embodiment, the generation unit 307 is specifically used for: the segmentation model includes: a convolutional neural network model, a converter segmentation network model, and a general segmentation model.

[0060] In one exemplary embodiment, the generation unit 307 is specifically used to: determine a generated bounding box hint, a point hint, and an edge hint based on the second candidate region; and input the generated bounding box hint, the point hint, the edge hint, and the fused defect model into the segmentation model to obtain a candidate defect mask.

[0061] In one exemplary embodiment, the third determining unit 308 is specifically used to: determine the second intersection-union ratio (IUU) between the candidate defect mask and the standard defect mask; obtain semantic similarity; perform weighted fusion of the second IUU and semantic similarity to obtain a joint scoring function; input each candidate defect mask into the joint scoring function to obtain the corresponding score value; and determine the corresponding candidate defect mask as the target defect when the score value is the largest.

[0062] Figure 4 This is a hardware block diagram of an electronic device provided according to an embodiment of the present disclosure. The electronic device 400 according to an embodiment of the present disclosure includes at least a processor and a memory for storing computer-readable instructions. When the computer-readable instructions are loaded and executed by the processor, the processor performs the steel plate surface defect detection method described in any of the preceding embodiments of the present disclosure.

[0063] Figure 4The illustrated electronic device 400 specifically includes a central processing unit (CPU) 401, a graphics processing unit (GPU) 402, and a memory 403. These units are interconnected via a bus 404. The CPU 401 and / or GPU 402 can function as the aforementioned processor, and the memory 403 can function as the aforementioned memory storing computer-readable instructions. Furthermore, the electronic device 400 may also include a communication unit 405, a storage unit 406, an output unit 407, an input unit 408, and an external device 409, all of which are also connected to the bus 404.

[0064] In summary, this disclosure provides a method, apparatus, and electronic device for detecting surface defects in steel plates. This disclosure utilizes different acquisition devices to simultaneously acquire multiple surface data points of the steel plate under test; these devices include industrial visible light cameras, infrared imaging devices, lidar devices, and spectral sensors. For any given surface data point, a first modal feature tensor is extracted and normalized to obtain a second modal feature tensor. The second modal feature tensor is used to indicate the first modal feature tensor with uniform dimension and scale. Based on a preset standard modal feature tensor and a first mapping relationship, first candidate defect regions for all second modal feature tensors are determined. The standard modal feature tensor is any one of the second modal feature tensors. The first mapping relationship includes a second mapping relationship between the second modal feature tensors and the steel plate under test, and multiple second modal feature tensors... The third mapping relationship between quantities; second candidate defect regions that satisfy the preset first intersection-union ratio are selected from all first candidate defect regions, and the cosine similarity of each second candidate defect region is determined; based on the cosine similarity and the fusion loss function, all second candidate defect regions are aligned to obtain third candidate defect regions; the fusion loss function is determined based on detection loss, masking loss and cross-modal consistency loss; all third candidate defect regions are weighted and fused to obtain fused defect regions; the fused defect regions and second candidate defect regions are input into the segmentation model to generate multiple candidate defect masks; the target defect is determined based on the candidate defect masks and the joint scoring function; the joint scoring function is determined based on the second intersection-union ratio and the semantic similarity between the candidate defect masks and the fused defect regions. In contrast to existing single-modal detection methods, which suffer from weak anti-interference capabilities, multi-modal detection methods that only involve simple global fusion, and a lack of region-level alignment and adaptive anti-interference mechanisms, this disclosure utilizes multi-device acquisition and achieves precise alignment of cross-modal features at the defect region granularity. It simultaneously constrains detection accuracy, segmentation accuracy, and cross-modal consistency through a fusion loss function. Furthermore, by employing a dual-intersection comparison and joint scoring mechanism, it effectively avoids noise mode interference, adapts to multi-scale defect detection needs, and improves defect identification sensitivity. In summary, the technical solution provided by this disclosure can improve the accuracy of steel surface defect detection and can be adapted to various application scenarios.

[0065] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.

[0066] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.

[0067] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.

[0068] Additionally, as used herein, the "or" used in a list of items beginning with "at least one" indicates a separate list, such that a list of, for example, "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not imply that the described example is preferred or better than other examples.

[0069] It should also be noted that in the systems and methods of this disclosure, the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalent solutions to this disclosure.

[0070] Various changes, substitutions, and modifications can be made to the technology described herein without departing from the teachings defined by the appended claims. Furthermore, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, events, means, methods, and actions described above. Currently existing or later-developed processes, machines, manufactures, events, means, methods, or actions that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Therefore, the appended claims include such processes, machines, manufactures, events, means, methods, or actions within their scope.

[0071] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.

[0072] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.

Claims

1. A method of detecting surface defects of a steel sheet, characterized by, The method includes: Multiple surface data of the steel plate under test are collected simultaneously using different acquisition devices; the different acquisition devices include: industrial visible light camera, infrared imaging equipment, lidar equipment and spectral sensor; For any of the surface data, the first modal feature tensor of the surface data is extracted and normalized to obtain the second modal feature tensor; the second modal feature tensor is used to indicate the first modal feature tensor with uniform dimension and scale. Based on a preset standard modal feature tensor and a first mapping relationship, first candidate defect regions for all second modal feature tensors are determined; the standard modal feature tensor is any one of the second modal feature tensors; the first mapping relationship includes: a second mapping relationship between the second modal feature tensor and the steel plate to be tested, and a third mapping relationship between multiple second modal feature tensors; Select second candidate defect regions that satisfy a preset first intersection-union ratio from all the first candidate defect regions, and determine the cosine similarity of each second candidate defect region; Based on the cosine similarity and the fusion loss function, all second candidate defect regions are aligned to obtain a third candidate defect region; the fusion loss function is determined based on detection loss, masking loss and cross-modal consistency loss. All the third candidate defect regions are weighted and fused to obtain the fused defect region; The fused defect region and the second candidate defect region are input into the segmentation model to generate multiple candidate defect masks; The target defect is determined based on the candidate defect mask and the joint scoring function; the joint scoring function is based on the second intersection-union ratio and the semantic similarity between the candidate defect mask and the fused defect region.

2. The method of claim 1, wherein, The target defects include at least one of the following: scratches, indentations, bubbles, microcracks, pinholes, and folds.

3. The method of claim 1, wherein, The determination of the first candidate defect regions for all second modal feature tensors based on the preset standard modal feature tensor and the first mapping relationship includes: Select any one of the second modal feature tensors as the standard modal feature tensor; Candidate defect boxes are determined based on the feature pyramid network. Based on the candidate defect box and the third mapping relationship, the candidate defect box is mapped to all second modality feature tensors except for the standard modality feature tensor; Based on the second mapping relationship, the mapped second modal feature tensor is associated with the position of the steel plate to be tested to obtain the first candidate defect region.

4. The method of claim 1, wherein, The step of selecting second candidate defect regions from all the first candidate defect regions that satisfy a preset first intersection-union ratio, and determining the cosine similarity of each second candidate defect region, includes: Determine the intersection-union ratio (IUU) between each of the first candidate defect regions and the standard defect mask; All first candidate defect regions whose cross-union ratio is greater than or equal to the first cross-union ratio are identified as second candidate defect regions; Based on the cosine similarity algorithm, the cosine similarity of any two second candidate defect regions is obtained.

5. The method of claim 1, wherein, The step of aligning all second candidate defect regions based on the cosine similarity and fusion loss function to obtain the third candidate defect region includes: For any two second candidate defect regions, the feature consistency error between the second candidate defect regions is determined based on the cosine similarity, and the cross-modal consistency loss is obtained. The detection loss is obtained; the detection loss is used to constrain the localization accuracy and category recognition accuracy of the second candidate defect region; The mask loss is obtained; the mask loss is used to constrain the contour segmentation accuracy of the second candidate defect region; The cross-modal consistency loss, the detection loss, and the masking loss are weighted and fused to obtain the fusion loss function; The third candidate defect region is obtained when the fusion loss value of all second candidate defect regions is minimized.

6. The method of claim 1, wherein, The segmentation models include: convolutional neural network models, converter segmentation network models, and general segmentation models.

7. The method of claim 1, wherein, The step of inputting the fused defect region and the second candidate defect region into the segmentation model to generate multiple candidate defect masks includes: Based on the second candidate region, generate box hints, dot hints, and edge hints; The generated box hints, the point hints, the edge hints, and the fused defect model are input into the segmentation model to generate multiple candidate defect masks.

8. The method of claim 1, wherein, The determination of the target defect based on the candidate defect mask and the joint scoring function includes: Determine the second intersection-union ratio (CUI) between the candidate defect mask and the standard defect mask; Obtain the semantic similarity; The second intersection-union ratio and the semantic similarity are weighted and fused to obtain the joint scoring function; Each of the candidate defect masks is input into the joint scoring function to obtain the corresponding score value; When the score value is the highest, the corresponding candidate defect mask is determined as the target defect.

9. A steel sheet surface defect detection device characterized by comprising: The device includes: The acquisition unit is used to simultaneously acquire multiple surface data of the steel plate under test using different acquisition devices; the different acquisition devices include: industrial visible light camera, infrared imaging device, lidar device and spectral sensor; An extraction unit is configured to extract a first modal feature tensor from any given surface data and perform normalization processing to obtain a second modal feature tensor; the second modal feature tensor is used to indicate the first modal feature tensor with uniform dimension and scale. The first determining unit is used to determine the first candidate defect region of all second modal feature tensors based on a preset standard modal feature tensor and a first mapping relationship; the standard modal feature tensor is any one of the second modal feature tensors; the first mapping relationship includes: a second mapping relationship between the second modal feature tensor and the steel plate to be tested, and a third mapping relationship between multiple second modal feature tensors; The second determining unit is used to filter second candidate defect regions that satisfy a preset first intersection-union ratio from all the first candidate defect regions, and to determine the cosine similarity of each second candidate defect region. An alignment unit is used to align all second candidate defect regions to obtain a third candidate defect region based on the cosine similarity and the fusion loss function; the fusion loss function is determined based on detection loss, masking loss and cross-modal consistency loss. A fusion unit is used to perform weighted fusion of all the third candidate defect regions to obtain a fused defect region; The generation unit is used to input the fused defect region and the second candidate defect region into the segmentation model to generate multiple candidate defect masks; The third determining unit is used to determine the target defect based on the candidate defect mask and the joint scoring function; the joint scoring function is based on the second intersection-union ratio and the semantic similarity between the candidate defect mask and the fused defect region.

10. An electronic device, comprising: include: Memory, used to store computer-readable instructions; as well as A processor for executing the computer-readable instructions, causing the electronic device to perform the method as described in any one of claims 1-8.