Steel coil end face defect detection method and system, training method and electronic equipment
By simultaneously acquiring and fusing RGB-D multimodal detection methods with two-dimensional images and three-dimensional point cloud data, the accuracy and robustness issues of steel coil end-face defect detection have been solved, enabling efficient identification and automated quality inspection of complex and subtle defects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 上海研视信息科技有限公司
- Filing Date
- 2026-03-17
- Publication Date
- 2026-04-17
AI Technical Summary
Existing steel coil end-face defect detection technologies suffer from low detection accuracy, weak robustness, and low degree of automation integration, making it particularly difficult to identify complex three-dimensional defects and micro-cracks.
By controlling the motion actuator to synchronously acquire two-dimensional image data and three-dimensional point cloud data, spatiotemporal registration is performed to generate RGB-D multimodal data. Surface texture and geometric structure features are extracted using a multimodal fusion neural network, and weighted fusion is performed through a cross-modal attention mechanism. Finally, the geometric feature consistency is verified by combining the three-dimensional point cloud data, and the defect detection results are output.
It significantly improves the accuracy and robustness of identifying complex and subtle defects, reduces the false detection rate, and achieves efficient automated steel coil end face quality inspection. It adapts to complex industrial site conditions and supports automatic marking and production line integration.
Smart Images

Figure CN121883480A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of quality inspection technology, and in particular to a method for detecting defects on the end face of steel coils, a training method for a defect detection model on the end face of steel coils, a defect detection system for the end face of steel coils, and electronic equipment. Background Technology
[0002] In the steel industry, steel coils are highly susceptible to various defects on their end faces during production, transportation, and handling, including cracks, scratches, dents, indentations, edge damage, and burrs. These defects not only affect the product's appearance but also directly interfere with subsequent processing steps (such as uncoiling, shearing, and stamping), and may even lead to equipment damage or product quality degradation, resulting in significant economic losses. Therefore, defect detection on the end faces of steel coils is a crucial step in ensuring production safety and product quality.
[0003] Traditional steel coil end-face defect detection mainly relies on manual visual inspection or two-dimensional image analysis based on fixed cameras. Manual inspection suffers from low efficiency, high labor intensity, susceptibility to subjective factors and fatigue, and difficulty in quantitatively assessing minute or three-dimensional defects. While traditional machine vision-based two-dimensional image analysis methods have achieved a degree of automation, their robustness heavily depends on lighting conditions and image quality. They are poorly resistant to interference from steel coil surface reflections, oxide scale, oil stains, etc., and cannot obtain three-dimensional geometric information such as defect depth and indentation amount, resulting in an extremely high rate of missed detection for typical three-dimensional defects such as "indentations," "protrusions," and "uneven thickness."
[0004] With technological advancements, the industry has begun to explore the introduction of deep learning and 3D sensing technologies. For instance, Chinese patent CN109829900A proposes a deep learning-based method for detecting defects on the end face of steel coils. This method trains a semantic segmentation network using synthesized grayscale images with defects to achieve pixel-level defect labeling. However, this approach relies solely on two-dimensional texture information and fails to adequately consider the three-dimensional geometric deformations present on the end face of the steel coil. For non-planar defects such as depressions and convexities, these often appear as blurry bright spots or shadows in the image, easily misidentified as reflective noise or oxide scale, resulting in a high false negative rate. More importantly, this method uses artificially synthesized data for training, making it difficult to generalize to the complex lighting and surface conditions in real industrial scenarios, thus lacking engineering practicality. Additionally, Chinese patent CN115294039A proposes a method for defect detection that fuses grayscale and depth maps. This method integrates two-dimensional and three-dimensional data through channel stitching and then inputs the data into a convolutional neural network (CSPNet) for processing. While this scheme incorporates 3D information, its fusion method remains at a simple data-level / channel-level stitching level. Essentially, it still uses a single network designed for images to process mixed data, failing to design a dedicated feature extraction and fusion mechanism for the heterogeneous characteristics of 2D images (texture, color) and 3D point clouds (geometry, spatial structure). Therefore, for complex defects requiring comprehensive utilization of texture appearance and 3D morphology information for accurate identification (such as distinguishing between fine cracks and scratches, and differentiating between shallow depressions and reflective spots), its detection accuracy and robustness still have significant room for improvement. Summary of the Invention
[0005] In view of the above-mentioned defects or deficiencies of the prior art, this application discloses a method, system and electronic equipment for detecting defects on the end face of steel coils, which can solve the problems of low detection accuracy, weak robustness and low degree of automation integration in the existing steel coil end face defect detection.
[0006] This application discloses a method for detecting defects on the end face of steel coils in its first aspect, comprising the following steps: The control motion actuator drives the two-dimensional image acquisition device and the three-dimensional point cloud acquisition device to dynamically scan the end face of the steel coil to be tested along a preset trajectory, and simultaneously acquire two-dimensional image data and three-dimensional point cloud data. The two-dimensional image data and the three-dimensional point cloud data are preprocessed and spatiotemporally registered to generate paired RGB-D multimodal data, wherein the pixels in the two-dimensional image and the points in the three-dimensional point cloud have a spatial correspondence. The RGB-D multimodal data is input into a pre-trained steel coil end-face defect detection model, and the following operations are performed: surface texture features and geometric structure features are extracted from the RGB-D multimodal data; the surface texture features and geometric structure features are dynamically weighted and fused using channel splicing and cross-modal attention mechanisms; based on the weighted fused features, the suspected defect results of the steel coil end face are output, wherein the suspected defect results represent at least one suspected defect and include the defect category, defect location, corresponding 3D depth information, and defect confidence level for each suspected defect; and Based on the suspected defect results, and in conjunction with the three-dimensional point cloud data, the geometric feature consistency of the suspected defects is verified to determine the defect detection result corresponding to the end face of the steel coil.
[0007] In some implementations of the first aspect, the control motion actuator drives the two-dimensional image acquisition device and the three-dimensional point cloud acquisition device to dynamically scan the end face of the steel coil to be tested along a preset trajectory, and simultaneously acquires two-dimensional image data and three-dimensional point cloud data, including: transporting the steel coil to be tested to the detection area, and triggering the detection process by the position detection component; controlling the motion actuator to drive the configured two-dimensional image acquisition device and the three-dimensional point cloud acquisition device to scan the end face of the steel coil along the preset trajectory, wherein the angle between the optical axis of the two-dimensional image acquisition device and the optical axis of the three-dimensional point cloud acquisition device and the normal of the end face of the steel coil is less than or equal to 30°; and through a hardware synchronization signal, enabling the two-dimensional image acquisition device and the three-dimensional point cloud acquisition device to simultaneously acquire two-dimensional image data and three-dimensional point cloud data during the scanning process, with a synchronization accuracy of less than or equal to 1 microsecond.
[0008] In some implementations of the first aspect, during the operation of controlling the motion actuator to drive the two-dimensional image acquisition device and the three-dimensional point cloud acquisition device to dynamically scan the end face of the steel coil to be tested along a preset trajectory: the preset trajectory is a spiral path or a grid path, and the scanning speed is 0.5 m / s to 2 m / s; the resolution of the two-dimensional image data is greater than or equal to 4096×4096 pixels; the point density of the three-dimensional point cloud data is greater than or equal to 100 points / square centimeter; wherein, when the diameter of the steel coil is greater than or equal to a preset diameter threshold and the flatness of the end face meets the preset conditions, the preset trajectory is planned using a spiral scanning path; when the diameter of the steel coil is less than the preset diameter threshold and the flatness of the end face does not meet the preset conditions, the preset trajectory is planned using a grid scanning path.
[0009] In some implementations of the first aspect, during the process of controlling the motion actuator to drive the two-dimensional image acquisition device and the three-dimensional point cloud acquisition device to dynamically scan the end face of the steel coil to be measured along a preset trajectory, an adaptive pose compensation step is also included: acquiring ranging information in real time and calculating the angle between the optical axis of either the two-dimensional image acquisition device or the three-dimensional point cloud acquisition device and the normal of the local area of the steel coil end face; dynamically correcting the end posture of the motion actuator based on the angle, so that the optical axes of the two-dimensional image acquisition device and the three-dimensional point cloud acquisition device are dynamically aligned with the normal direction of the local area of the steel coil end face.
[0010] In some implementations of the first aspect, preprocessing the two-dimensional image data includes: performing adaptive contrast enhancement processing on the two-dimensional image data, using a filtering algorithm to remove noise, and extracting the outline of the steel coil end face through edge detection to crop invalid areas; preprocessing the three-dimensional point cloud data includes: removing outliers through statistical filtering, reducing point cloud density through downsampling, and then smoothing the downsampled point cloud; spatiotemporal registration of the two-dimensional image data and the three-dimensional point cloud data includes: projecting the three-dimensional point cloud data onto the two-dimensional image coordinate system using calibration parameters to generate RGB-D multimodal data with depth information.
[0011] In some implementations of the first aspect, the dynamic weighted fusion of the surface texture features and the geometric structure features through channel splicing and cross-modal attention mechanisms includes: splicing the surface texture features and the geometric structure features in the channel dimension; using the spatial saliency weights from the surface texture features to guide the allocation of aggregate weights from the geometric structure features to enhance geometric distortion features in low-contrast regions; and using the depth variation weights from the geometric structure features to enhance the texture response of suspected regions in the surface texture features to filter out non-physical depth interference in complex texture backgrounds.
[0012] In some implementations of the first aspect, the steel coil end-face defect detection model is formed based on training of a multimodal fusion neural network, which includes: a two-dimensional feature extraction branch, employing a deep convolutional neural network with embedded spatial attention mechanism, for extracting surface texture features from the two-dimensional image data; a three-dimensional feature extraction branch, employing a point cloud processing network with hierarchical feature aggregation, for extracting geometric structure features from the three-dimensional point cloud data; and a fusion decision layer, for concatenating the surface texture features and the geometric structure features in the channel dimension, and dynamically weighting and fusing the concatenated features through a cross-modal attention mechanism, and then simultaneously performing defect category classification, defect location regression, three-dimensional depth information calculation, and defect confidence assessment based on the weighted fused features to generate a suspected defect result corresponding to the steel coil end-face.
[0013] In some implementations of the first aspect, the geometric feature consistency verification of the suspected defect includes: extracting a corresponding local point cloud subset from the three-dimensional point cloud data based on the defect location of the suspected defect; performing plane fitting on the background points in the local point cloud subset using a random sampling consensus algorithm to construct a local reference plane reflecting the local pose state of the current steel coil end face; calculating the normal distance of each point in the local point cloud subset relative to the local reference plane, and extracting the maximum depth value and the average depth difference; if the maximum depth value or the average depth difference is greater than a preset physical defect threshold, then the suspected defect is determined to be a real defect.
[0014] In some implementations of the first aspect, after determining the defect detection result corresponding to the end face of the steel coil, the method further includes: sending the defect detection result to the production line control system, whereby the production line control system generates process control instructions based on the defect detection result to perform any one or more operations among automatic marking, sorting, and production line alarm.
[0015] This application discloses a training method for a steel coil end-face defect detection model in its second aspect, comprising the following steps: Multiple steel coil end face samples are acquired. Each steel coil end face sample contains synchronously acquired two-dimensional image data and three-dimensional point cloud data. Defect truth information is labeled for each steel coil end face sample. The defect truth information includes defect category, corresponding defect area in two-dimensional image, and corresponding geometric deformation area in three-dimensional point cloud. The two-dimensional image data and the three-dimensional point cloud data in each steel coil end face sample are preprocessed and spatiotemporally registered to form paired RGB-D samples; by performing the above processing on multiple steel coil end face samples, a training sample set is generated. The training sample set is input into a multimodal fusion neural network for forward propagation, and the corresponding defect prediction result is output. The multimodal fusion neural network is configured to extract surface texture features from the two-dimensional image data and geometric structure features from the three-dimensional point cloud data, and to dynamically weight and fuse the surface texture features and the geometric structure features through channel splicing and cross-modal attention mechanisms. Calculate the loss between the defect prediction result and the defect ground truth information; the loss includes defect classification loss, defect localization regression loss, and geometric consistency loss, wherein the geometric consistency loss is used to constrain the overlap between the two-dimensional defect region and the three-dimensional depth anomaly region predicted by the model in spatial projection; and Based on the aforementioned loss, the backpropagation algorithm is used to update the network parameters of the multimodal fusion neural network, and the optimization is iteratively performed until the model converges, thus obtaining a trained steel coil end face defect detection model.
[0016] In some implementations of the second aspect, the training method of the steel coil end-face defect detection model further includes: constructing a multimodal fusion neural network, wherein the multimodal fusion neural network includes a two-dimensional feature extraction branch, a three-dimensional feature extraction branch, and a fusion decision layer; wherein the two-dimensional feature extraction branch adopts a deep convolutional neural network with embedded spatial attention mechanism to extract surface texture features from the two-dimensional image data; the three-dimensional feature extraction branch adopts a point cloud processing network with hierarchical feature aggregation to extract geometric structure features from the three-dimensional point cloud data; the fusion decision layer is used to concatenate the surface texture features and the geometric structure features in the channel dimension, and to perform dynamic weighted fusion based on the concatenated features through a cross-modal attention mechanism.
[0017] This application discloses a steel coil end face defect detection system in a third aspect, comprising: The data acquisition module includes a motion actuator and a two-dimensional image acquisition device and a three-dimensional point cloud acquisition device mounted on the motion actuator. The motion actuator drives the two-dimensional image acquisition device and the three-dimensional point cloud acquisition device to dynamically scan the end face of the steel coil to be tested along a preset trajectory, and simultaneously acquire two-dimensional image data and three-dimensional point cloud data. The data processing module is used to preprocess and spatiotemporally register the two-dimensional image data and the three-dimensional point cloud data to generate paired RGB-D multimodal data, wherein the pixels in the two-dimensional image and the points in the three-dimensional point cloud have a spatial correspondence. The data analysis module is configured with a steel coil end-face defect detection model. This model receives RGB-D multimodal data and extracts surface texture and geometric features from it. The surface texture and geometric features are dynamically weighted and fused using channel splicing and cross-modal attention mechanisms. Based on the weighted fused features, the module outputs suspected defect results for the steel coil end-face. Each suspected defect result represents at least one suspected defect and includes the defect category, defect location, corresponding 3D depth information, and defect confidence level for each suspected defect. The verification module is used to perform geometric feature consistency verification on the suspected defect based on the suspected defect results and the three-dimensional point cloud data, and to determine the defect detection result corresponding to the end face of the steel coil.
[0018] In a fourth aspect, this application discloses an electronic device including a processor and a memory, wherein the memory stores a computer program that, when executed by the processor, implements the steel coil end-face defect detection method as described above.
[0019] In its fifth aspect, this application discloses a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steel coil end-face defect detection method as described above.
[0020] Compared with the prior art, the steel coil end-face defect detection method, steel coil end-face defect detection system, electronic equipment, and computer-readable storage medium disclosed in this application have at least the following beneficial effects: This application constructs RGB-D multimodal data with pixel-level correspondence by spatiotemporally registering two-dimensional image data and three-dimensional point cloud data. In the steel coil end-face defect detection model, surface texture features and geometric structure features are jointly modeled and fused for analysis. This enables the steel coil end-face defect detection model to simultaneously utilize surface texture variations and spatial geometric undulation information for defect identification. Compared to detection methods based solely on two-dimensional images, this deep fusion of multimodal information can more comprehensively characterize the features of steel coil end-face defects, improving the ability to identify low-contrast defects and complex-shaped defects.
[0021] This application innovatively designs a steel coil end-face defect detection model. This model can extract surface texture features and geometric structure features from RGB-D multimodal data, and then use channel splicing and cross-modal attention mechanisms to dynamically weight and fuse the surface texture features and geometric structure features. This achieves deep complementarity and synergy of the two types of heterogeneous information at the feature level, improves the speed of steel coil end-face defect detection, and more importantly, significantly improves the accuracy and robustness of identifying complex and subtle defects (such as micro-cracks and dents under reflective interference).
[0022] This application introduces a geometric feature consistency verification mechanism based on 3D point cloud data after the suspected defect results of the steel coil end face defect detection model. By analyzing the geometric features of the suspected defect area, it determines whether it meets the judgment conditions of real physical defects, thereby re-verifying the output results of the steel coil end face defect detection model. This can effectively filter out false defects caused by factors such as surface texture changes and lighting interference, reduce the false detection rate, and improve the reliability of the detection results.
[0023] This application acquires steel coil end face data through dynamic scanning and combines adaptive attitude compensation and high-precision spatiotemporal registration processing, which can effectively adapt to complex industrial site conditions such as changes in steel coil size, uneven end face, and equipment posture deviation, thereby improving the stability and robustness of data acquisition in actual production environments.
[0024] The defect detection results output by this application have clear defect categories, locations, and physical characteristics. They can directly interact with the production line control system to realize various production line operations such as automatic marking, sorting, and alarming of defective steel coils, forming a closed-loop process that combines detection and production control, thereby improving the automation level and industrial application value of steel coil end face quality inspection. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 The diagram shown is a flowchart of one embodiment of the steel coil end-face defect detection method of this application.
[0027] Figure 2 The diagram shows a six-axis industrial robot driving a two-dimensional image acquisition device and a three-dimensional point cloud acquisition device to collect data from the end face of a steel coil.
[0028] Figure 3 The diagram shows the architecture of a multimodal fusion neural network in one embodiment.
[0029] Figure 4 The diagram shown is a flowchart of one embodiment of the steel coil end-face defect detection method of this application.
[0030] Figure 5 The diagram shown is a flowchart of one embodiment of the steel coil end-face defect detection model training method of this application.
[0031] Figure 6 The diagram shown is a structural schematic of one embodiment of the steel coil end-face defect detection system of this application.
[0032] Figure 7 The diagram shown is a structural schematic of an electronic device according to an embodiment of this application. Detailed Implementation
[0033] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, unless otherwise specified, the following embodiments and features in the embodiments can be combined with each other.
[0034] It should be noted that the illustrations disclosed in the following embodiments are only schematic representations of the basic concept of this application. Therefore, the drawings only show the components related to this application and are not drawn according to the number, shape and size of the components in actual implementation. In actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0035] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the present application. However, it will be apparent to those skilled in the art that embodiments of the present application may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the present application.
[0036] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this application described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion.
[0037] Unless otherwise stated, the term "multiple" means two or more. In the embodiments of this application, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B means: A or B. The term "and / or" describes an association relationship between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or, A and B.
[0038] Generally, traditional steel coil end-face defect detection mainly relies on manual visual inspection or two-dimensional image analysis methods based on single 2D vision, which suffers from poor anti-interference ability and high false detection rate. Existing technical solutions that attempt to incorporate 3D data mostly remain at the stage of simple "data stitching" or "result filtering," essentially still using a single network designed for images to process mixed data, resulting in insufficient detection accuracy and robustness. Furthermore, existing defect detection systems cannot effectively distinguish between visual noise and real physical defects, leading to a persistently high false detection rate, and there is a lack of subsequent verification mechanisms for preliminary detection results.
[0039] In view of this, this application discloses a method for detecting defects on the end face of a steel coil. By controlling a motion actuator to drive a two-dimensional image acquisition device and a three-dimensional point cloud acquisition device to simultaneously scan the end face of the steel coil, paired two-dimensional image data and three-dimensional point cloud data are obtained. After preprocessing and spatiotemporal registration, RGB-D multimodal data corresponding to pixels and three-dimensional points are generated. This RGB-D multimodal data is input into a steel coil end face defect detection model, which extracts two-dimensional texture and geometric structure features respectively. After channel splicing and cross-modal attention mechanism fusion, a suspected defect containing category, location and confidence level is output. Finally, the suspected defect is geometrically verified based on the three-dimensional point cloud to analyze the actual depth and shape, eliminate false defects, and output the final detection result. This end-face defect detection method realizes a closed-loop detection of "dynamic acquisition-multimodal fusion-three-dimensional verification", effectively overcoming interference such as surface reflection and complex shape, and significantly improving the accuracy and robustness of identifying two-dimensional and three-dimensional composite defects such as micro-cracks and dents. It can be seamlessly integrated into existing production lines to meet the industrial site's demand for high-precision and high-efficiency end-face quality inspection, and has significant engineering application value and promotion prospects.
[0040] Please see Figure 1 The diagram shows a flowchart of one embodiment of the steel coil end-face defect detection method of this application.
[0041] like Figure 1 As shown, the method for detecting defects on the end face of steel coils in the embodiment includes the following steps: Step S101: Control the motion actuator to drive the two-dimensional image acquisition device and the three-dimensional point cloud acquisition device to dynamically scan the end face of the steel coil to be tested along the preset trajectory, and simultaneously acquire two-dimensional image data and three-dimensional point cloud data.
[0042] Step S103: Preprocess and spatiotemporally register the two-dimensional image data and the three-dimensional point cloud data to generate paired RGB-D multimodal data.
[0043] Step S105: Input the RGB-D multimodal data into the trained steel coil end face defect detection model and output the suspected defect results of the steel coil end face.
[0044] Step S107: Based on the suspected defect results and combined with the three-dimensional point cloud data, perform geometric feature consistency verification on the suspected defects to determine the defect detection results corresponding to the end face of the steel coil.
[0045] This application acquires two-dimensional image data and three-dimensional point cloud data of the end face of steel coils through dynamic scanning, and forms RGB-D multimodal data after spatiotemporal registration. This enables the steel coil end face defect detection model to identify defects based on RGB-D multimodal data, output suspected defect results, and then combine the three-dimensional point cloud data to verify the geometric feature consistency of suspected defect areas to obtain the final defect detection result. Compared with the existing technology, this improves the speed of steel coil end face defect detection and significantly enhances the accuracy and robustness of identifying complex and subtle defects.
[0046] The following provides a detailed explanation of each of the above steps.
[0047] Step S101: Control the motion actuator to drive the two-dimensional image acquisition device and the three-dimensional point cloud acquisition device to dynamically scan the end face of the steel coil to be tested along the preset trajectory, and simultaneously acquire two-dimensional image data and three-dimensional point cloud data.
[0048] In some embodiments, in step S101, controlling the motion actuator to drive the two-dimensional image acquisition device and the three-dimensional point cloud acquisition device to dynamically scan the end face of the steel coil to be tested along a preset trajectory, and simultaneously acquiring two-dimensional image data and three-dimensional point cloud data, may further include: First, the steel coil to be tested is conveyed to the testing area, and the testing process is triggered by a position detection component. The position detection component is used to detect whether the steel coil to be tested has reached the designated position. When the steel coil to be tested reaches the designated position, the position detection component can send a trigger signal to initiate the testing process.
[0049] For example, the position detection component may be a photoelectric sensor, a non-contact detection device that uses the emission and reception of light beams to detect the presence, position, distance, color, or surface condition of an object. Generally, a photoelectric sensor may include a light source, a light receiver, and a signal processor. In the detection of defects on the end face of steel coils, the photoelectric sensor can be installed at the detection station to accurately detect whether the steel coil has been delivered to the detection station, and use this as the starting signal to trigger the entire detection process.
[0050] For example, the position detection component may be a laser rangefinder, which calculates the distance to the target by generating a laser beam and measuring its round-trip time or phase difference. In the detection of defects on the end face of steel coils, the laser rangefinder can be placed at the inspection station to accurately detect whether the steel coil has been delivered to the inspection station, and use this as the starting signal to trigger the entire inspection process.
[0051] For example, the position detection component may be a visual positioning component, such as an industrial camera combined with image recognition. In the detection of defects on the end face of steel coils, an industrial camera is set at the inspection station to capture images of the steel coils. Image processing algorithms (such as edge detection, template matching, deep learning, etc.) are used to identify the position, contour, and orientation of the coils, determine whether the steel coils have been delivered to the inspection station, and use this as the starting signal to trigger the entire inspection process.
[0052] Next, the control motion actuator drives the configured two-dimensional image acquisition device and three-dimensional point cloud acquisition device to scan the end face of the steel coil along a preset trajectory.
[0053] In some embodiments, the motion actuator may be, for example, a six-axis industrial robot, the two-dimensional image acquisition device may be, for example, a high-resolution industrial line scan camera, and the three-dimensional point cloud acquisition device may be, for example, a high-precision 3D laser line scan camera. See also... Figure 2 This is a schematic diagram showing the state of a six-axis industrial robot driving a two-dimensional image acquisition device and a three-dimensional point cloud acquisition device to collect data from the end face of a steel coil. (Example:) Figure 2 As shown, both the industrial line scan camera and the 3D laser line scan camera are located at the end of the six-axis industrial robot 21. By controlling the six-axis industrial robot 21, the two-dimensional image acquisition device and the three-dimensional point cloud acquisition device can be driven to dynamically scan the end face 31 of the steel coil 30 to be tested along a preset trajectory, and simultaneously acquire two-dimensional image data and three-dimensional point cloud data.
[0054] Multimodal data of the steel coil end face is acquired through the collaborative operation of an industrial line scan camera and a 3D laser line scan camera. The industrial line scan camera continuously acquires high-resolution line images during high-speed movement and stitches them together to generate a complete two-dimensional surface texture image. Its output image resolution is no less than 4096×4096 pixels, sufficient to clearly present the texture features of micron-level cracks, scratches, and other minute defects. Simultaneously, the 3D laser line scan camera acquires the three-dimensional morphology information of the steel coil end face in real time using laser triangulation, generating dense and accurate point cloud data with a point density of no less than 100 points / square centimeter. This effectively characterizes the depth and contour of geometric defects such as depressions, protrusions, and pyramidal shapes. These resolution and point density parameters were determined after extensive experimental verification: excessively low image resolution leads to the loss of subtle texture features, while insufficient point cloud density makes it difficult to accurately reconstruct local surface morphology, thus affecting the defect detection rate and quantitative accuracy. Therefore, by configuring acquisition equipment that meets the above performance requirements, the system can provide high-quality, high signal-to-noise ratio basic data support for subsequent multimodal fusion analysis and three-dimensional geometric verification while ensuring detection efficiency.
[0055] In an embodiment of the invention, a motion actuator (e.g., a six-axis industrial robot) drives an industrial line scan camera and a 3D laser line scan camera integrated at its end to perform a full-coverage dynamic scan of the steel coil end face along a preset trajectory. To ensure data acquisition quality, the mounting posture of the industrial line scan camera and the 3D laser line scan camera is precisely calibrated to ensure that the angle between their optical axes and the local normal of the steel coil end face does not exceed 30°. This angle limitation effectively avoids problems such as image perspective distortion, edge blurring, and laser reflection signal attenuation caused by excessively large viewing angles. Especially in the outer edge region of the steel coil, it can significantly improve the clarity of two-dimensional textures and the integrity of three-dimensional point clouds, laying the foundation for subsequent high-precision defect identification.
[0056] Meanwhile, to achieve strict alignment of multimodal data in the spatiotemporal dimensions, the system employs a hardware-level synchronization mechanism. In some embodiments, a dedicated synchronization trigger signal (e.g., synchronization I / O or hardware pulse) controls the industrial line scan camera and the 3D laser line scan camera to work together, ensuring that each line of two-dimensional image and its corresponding three-dimensional point cloud strip are precisely matched at the time of acquisition, with a synchronization accuracy of less than or equal to 1 microsecond.
[0057] Furthermore, the motion actuator drives the industrial line scan camera and the 3D laser line scan camera to perform a full-coverage dynamic scan of the steel coil end face along a preset trajectory. The scanning trajectory is not fixed but can be adaptively planned according to the geometric characteristics of the steel coil. In some cases, when the steel coil diameter is greater than or equal to a preset diameter threshold (e.g., 1.0 meter, 1.5 meters) and the end face flatness deviation is less than a preset condition (e.g., 2 mm), the planned preset trajectory adopts a spiral scanning path, continuously expanding outward from the inner hole. This has the advantages of continuous path, high efficiency, and uniform coverage, and is suitable for rapid detection of large-area flat end faces. In some embodiments, when the steel coil diameter is less than the preset diameter threshold (e.g., 1.0 meter, 1.5 meters) or when there are irregular deformations such as local collapse or warping on the end face, the planned preset trajectory adopts a grid scanning path, which consists of a series of parallel straight line segments. This can flexibly adapt to complex surface undulations and ensure that sufficient density and accuracy of point cloud data can still be obtained in geometrically abnormal areas. This adaptive trajectory planning strategy based on operational condition awareness balances detection efficiency and data fidelity, and is a key technological element for achieving highly robust and high-coverage online quality inspection. The entire scanning process operates at a speed between 0.5 m / s and 2 m / s, balancing detection efficiency and data quality. The entire process is unmanned, enabling real-time processing of detection data and decision feedback, seamlessly integrating with PLCs and smart factory systems, and driving the digital transformation of steel production quality control.
[0058] Furthermore, to address the practical situation where the end face of the steel coil is not an ideal plane, an adaptive pose compensation mechanism is introduced. During the process of controlling the motion actuator to drive the two-dimensional image acquisition device (e.g., an industrial line scan camera) and the three-dimensional point cloud acquisition device (e.g., a 3D laser line scan camera) to dynamically scan the end face of the steel coil to be measured along a preset trajectory, distance measurement information (e.g., local point cloud data from the 3D laser camera) is acquired in real time. Based on this, the angle between the optical axis of either the industrial line scan camera or the 3D laser line scan camera and the normal of the local area of the steel coil end face is calculated (i.e., the angle between the optical axis of the industrial line scan camera and the normal of the local area of the steel coil end face, or the angle between the optical axis of the 3D laser line scan camera and the normal of the local area of the steel coil end face). Based on this angle, the end-effector posture of the motion actuator (e.g., a six-axis industrial robot) is dynamically corrected. For example, when this angle exceeds a preset tolerance (e.g., 5°), the end-effector posture of the six-axis industrial robot is dynamically corrected so that the optical axes of the industrial line scan camera and the 3D laser line scan camera are always perpendicular to the local tangent plane of the steel coil end face. This adaptive pose compensation mechanism effectively solves the imaging viewpoint shift problem caused by the step-like offset or local collapse formed by end face warping and interlayer misalignment. It significantly alleviates the resulting phenomena such as 2D image blurring, edge distortion, and sparse or missing 3D point clouds, ensuring that high-resolution texture images and high-density, high signal-to-noise ratio point cloud data can still be stably acquired in areas with complex geometric deformation, providing a reliable data foundation for subsequent multimodal fusion and geometric verification.
[0059] This application integrates six-axis industrial robot motion control with multi-sensor collaborative scanning to achieve blind-spot-free coverage of the steel coil end face, supports the inspection of full-size steel coils (outer diameter 800 mm to 2500 mm), and the inspection time for a single coil can be less than 30 seconds (including scanning and analysis). It supports real-time online inspection at production line speeds greater than or equal to 2 m / s, meeting the needs of high-speed industrial production.
[0060] Through step S101, two-dimensional image acquisition equipment and three-dimensional point cloud acquisition equipment can be used to dynamically scan the end face of the steel coil to be tested along a preset trajectory and then simultaneously acquire two-dimensional image data and three-dimensional point cloud data.
[0061] Step S103: Preprocess and spatiotemporally register the two-dimensional image data and the three-dimensional point cloud data to generate paired RGB-D multimodal data.
[0062] After completing the dynamic scan to acquire data about the end face of the steel coil, the acquired two-dimensional image data and three-dimensional point cloud data are preprocessed.
[0063] Regarding the two-dimensional image data, considering that the steel coil end face may experience uneven lighting, surface reflection, and dust adhesion in actual production environments, quality optimization processing of the two-dimensional image data is necessary. Specifically, the preprocessing of the two-dimensional image data includes: adaptive contrast enhancement processing to strengthen the grayscale difference between the defect area and the background area; and suppressing image noise through filtering algorithms to prevent noise features from interfering with the subsequent recognition process. Based on this, the outline of the steel coil end face is extracted through edge detection, and the image is cropped according to the edge outline information of the steel coil end face, retaining only the effective area related to the steel coil end face, thereby reducing the impact of the background area on the detection results.
[0064] The adaptive contrast enhancement processing may include, but is not limited to, methods such as Adaptive Histogram Equalization (AHE) and Contrast Limited Adaptive Histogram Equalization (CLAHE). Taking Adaptive Histogram Equalization (AHE) as an example, to address the problem of large differences in illumination between different areas (such as the inner and outer rings) of the steel coil end face, Adaptive Histogram Equalization divides the image into multiple small blocks and performs histogram equalization independently on each block, thereby significantly enhancing local contrast and making tiny cracks (with a width as low as 0.1 mm) clearly visible in the image.
[0065] The filtering algorithm can use median filtering to effectively remove noise caused by sensor thermal noise or dust, while better preserving the edge information of the image and avoiding edge blurring caused by Gaussian filtering.
[0066] The edge detection and cropping can employ Canny edge detection and ROI cropping. The Canny operator accurately extracts the circular outer contour and inner hole contour of the steel coil, using this as a basis to crop away the background area in the image, forming a precise region of interest (ROI). This not only significantly reduces the amount of data required for subsequent processing, but more importantly, it eliminates the negative impact of background interference on the defect detection model.
[0067] 3D point cloud data also contains a significant amount of noise and redundant information. Due to potential variations in material reflection characteristics or environmental interference during the 3D laser line scanning process, a small number of outliers may exist in the point cloud data. Directly including these outliers in geometric analysis can easily lead to misjudgments. Therefore, preprocessing of 3D point cloud data includes denoising through statistical filtering to improve the overall reliability of the data. Following this, the system downsamples the point cloud data to reduce its density, and then smooths the downsampled point cloud to optimize its surface smoothness. This reduces the point cloud density while preserving the main geometric features, thereby improving subsequent computational efficiency and enhancing the stability of geometric analysis.
[0068] The point cloud data is denoised to remove outliers caused by strong reflections, splashes, or measurement errors. This denoising process can employ statistical filtering methods, identifying and removing outliers based on the geometric distribution characteristics of local neighborhoods. For example, the statistical filtering uses the Statistical Outlier Removal (SOR) algorithm. This algorithm calculates the average distance between each point and its nearest neighbors, calculates the mean μ and standard deviation σ of the average distances of all points, and removes points with an average distance greater than μ + λσ (λ is a threshold, typically 1.0).
[0069] The original point cloud data is downsampled to reduce data redundancy and improve subsequent processing efficiency. For example, the downsampling can be performed using a voxel grid downsampling (VGD) method, which divides the three-dimensional space into cubic units (i.e., voxels) with a side length of 0.5 mm. Only one representative point (preferably the centroid of all points within that voxel) is retained in each non-empty voxel, thereby reducing the amount of point cloud data by about 50% with almost no loss of key geometric details.
[0070] The downsampled point cloud is smoothed to suppress high-frequency noise and enhance the discernibility of macroscopic geometric features. For example, the smoothing process can employ Moving Least Squares (MLS), which fits a smooth surface to the local neighborhood of the point cloud and projects the original points onto the reconstructed surface. This effectively filters out high-frequency noise caused by sensor errors or surface perturbations, making structural defects or contour features such as depressions and protrusions clearer and more prominent.
[0071] To achieve collaborative analysis of 2D images and 3D point clouds in a unified space, high-precision spatiotemporal registration is performed on the acquired 2D image data and 3D point cloud data. This spatiotemporal registration includes projecting the 3D point cloud data onto the 2D image coordinate system using calibration parameters to generate RGB-D multimodal data. Specifically, The spatiotemporal registration includes two stages: offline calibration and online registration. During the offline calibration phase, a high-precision calibration board (e.g., a checkerboard pattern) is used in the initial deployment stage to calibrate the industrial line scan camera based on a specific calibration method (e.g., the Zhang Zhengyou calibration method). This process obtains its intrinsic parameter matrix K and lens distortion coefficients (including radial and tangential distortion coefficients). Simultaneously, the extrinsic parameters (i.e., rotation matrix R and translation vector T) between the 3D laser line scan camera and the industrial line scan camera are calibrated. This calibration process establishes the precise geometric relationship between the industrial line scan camera and the 3D laser line scan camera in the world coordinate system.
[0072] During the online registration stage, in each scan, the original two-dimensional image is first subjected to distortion correction, and the nonlinear distortion caused by the lens is eliminated using the distortion coefficients obtained from the calibration above. Subsequently, based on the corrected image coordinate system, each three-dimensional space point (Xw, Yw, Zw) is mapped to the two-dimensional image coordinate system according to the following projection model using the intrinsic parameter matrix K and the extrinsic parameter [R|T]: This projection mapping assigns a corresponding depth value Z to each valid pixel (u, v) in a 2D image, thereby constructing a spatially aligned RGB-D image and generating RGB-D multimodal data that simultaneously contains color and depth information. The RGB channels retain the original 2D texture information, while the depth channel provides precise 3D geometric information. Thus, a one-to-one correspondence is established between 2D pixels and 3D spatial points, achieving depth fusion of multimodal data within a unified coordinate framework.
[0073] In step S103, the original, noisy heterogeneous data (two-dimensional image data and three-dimensional point cloud data) are transformed into clean, concise RGB-D multimodal data with strictly aligned pixels and points.
[0074] Step S105: Input the RGB-D multimodal data into the trained steel coil end face defect detection model and output the suspected defect results of the steel coil end face.
[0075] To achieve high-precision and robust detection of complex defects on the end face of steel coils (such as micro-cracks, three-dimensional depressions, edge damage, etc.), this application constructs a defect detection model for the end face of steel coils.
[0076] By inputting the RGB-D multimodal data into the pre-trained steel coil end-face defect detection model, the following operations can be performed: extracting surface texture features and geometric structure features from the RGB-D multimodal data; dynamically weighting and fusing the surface texture features and geometric structure features through channel splicing and cross-modal attention mechanisms; and outputting the suspected defect results of the steel coil end face based on the weighted fused features.
[0077] The extraction of surface texture features and geometric structure features from the RGB-D multimodal data includes extracting surface texture features from the two-dimensional image data and extracting geometric structure features from the three-dimensional point cloud data.
[0078] The surface texture features and geometric structure features are dynamically weighted and fused using channel splicing and cross-modal attention mechanisms, including at least the following: The surface texture features and the geometric structure features are spliced together in the channel dimension.
[0079] The spatial saliency weights from the surface texture features are used to guide the aggregation weight allocation from the geometric structure features to enhance the geometric distortion features in low-contrast regions.
[0080] By utilizing the depth variation weights of the geometric structural features, the texture response of suspected regions in the surface texture features is enhanced to filter out non-physical depth interference in complex texture backgrounds.
[0081] The step of outputting the suspected defect results of the steel coil end face based on the weighted fusion features includes: simultaneously performing defect category classification, defect location regression, three-dimensional depth information calculation and defect confidence assessment based on the weighted fusion features, and generating suspected defect results corresponding to the steel coil end face.
[0082] In some embodiments, the steel coil end-face defect detection model is a multimodal fusion neural network.
[0083] Please see Figure 3 The diagram shows the architecture of a multimodal fusion neural network in one embodiment.
[0084] Combination Figure 3 As shown, the multimodal fusion neural network includes: a two-dimensional feature extraction branch, a three-dimensional feature extraction branch, and a fusion decision layer.
[0085] The two-dimensional feature extraction branch employs a deep convolutional neural network with embedded spatial attention mechanism to extract surface texture features from the two-dimensional image data. The processing flow of this branch is shown in Figure 3. The overall architecture consists of an improved ResNet-50 backbone network, and a spatial attention mechanism (SAM) is introduced into the key residual blocks to enhance the perception of defective regions.
[0086] First, the input 2D image enters the improved ResNet-50 backbone network, which extracts visual features from the image step by step through multiple convolutional operations: from low-level basic texture information such as edges and lines, to high-level semantic structures (such as crack direction, indentation contours, and oxide distribution). Because steel coil end-face images generally suffer from uneven illumination, reflective interference, and weak local details, traditional convolutional networks are easily affected by background noise, causing defect features to be obscured.
[0087] To address this, a spatial attention mechanism is embedded in multiple residual blocks of the network (as shown in the diamond-shaped decision boxes in the figure). This mechanism dynamically enhances the response of defect-related regions by adaptively weighting the spatial location of the feature maps. Specifically, the spatial attention module calculates the importance weight of each spatial location, significantly improving the feature representation ability of elongated gray-scale abrupt regions (such as cracks) and regions with strong light-dark contrast (such as recessed edges), while effectively suppressing interference signals from uniform backgrounds (such as oxide scale and normal reflection).
[0088] Finally, the attention-enhanced features are processed by global pooling and output as a 512-dimensional texture feature vector. This vector accurately encodes the texture structure and defect appearance information of the steel coil end face surface, and has good discriminativeness and robustness. It can serve as an important basis for subsequent defect identification and geometric consistency verification.
[0089] The three-dimensional feature extraction branch employs a hierarchical feature aggregation point cloud processing network to extract geometric structural features from the three-dimensional point cloud data.
[0090] In some embodiments, the 3D feature extraction branch employs a point cloud processing network based on the PointNet++ algorithm (also known as the PointNet++ point cloud network) to specifically process high-quality 3D point cloud data after preprocessing (such as denoising, downsampling, and smoothing). The processing flow of this branch is shown in Figure 3. The overall architecture consists of multiple hierarchical modules that sequentially complete geometric feature extraction from local to global.
[0091] First, the input 3D point cloud data enters the PointNet++ point cloud network, which gradually constructs a multi-scale point cloud representation through a hierarchical "sampling-grouping-feature aggregation" mechanism. In each layer, the network samples the current point set to generate keypoints, then divides local neighborhoods around these keypoints, and performs feature extraction operations within each neighborhood.
[0092] Next, in the component local feature aggregation stage, the network encodes the features of each point in the local neighborhood, capturing the relative positional relationship between the points and local geometric properties (such as normal vector, curvature, gradient, etc.), forming local geometric features with spatial semantics.
[0093] The process then moves to the local geometric feature capture stage, which further refines the geometric structure information of the local area, enhances the sensitivity to minor deformations (such as depressions, cracks, and edge lifting), and improves the feature expression capability.
[0094] Based on this, the system uses a global morphological feature extraction module to upsample and fuse all local features step by step, and finally generates a high-dimensional geometric feature vector that can comprehensively characterize the macroscopic morphology (such as overall flatness and curvature distribution) and microscopic defects (such as indentation depth and crack orientation) of the steel coil end face.
[0095] The final output is a 256-dimensional geometric feature vector, which serves as the input for subsequent defect classification and geometric consistency verification. This feature vector not only possesses good discriminative ability but also strong robustness, effectively handling noise interference and deformation complexity in real-world industrial scenarios.
[0096] The two-dimensional feature extraction branch and the three-dimensional feature extraction branch described above depict the state of the steel coil end face from different dimensions.
[0097] The fusion decision layer is used to perform channel splicing and utilizes the cross-modal attention mechanism (CMA) to fuse surface texture features from the two-dimensional feature extraction branch and geometric structure features from the three-dimensional feature extraction branch. Based on the weighted fused features, it performs defect category classification, localization regression, and confidence assessment, and finally generates a comprehensive suspected defect result corresponding to the end face of the steel coil.
[0098] The processing flow of the fusion decision layer is shown in Figure 3. The overall architecture consists of three key stages: basic feature integration, cross-modal dynamic weighting, and multi-task joint reasoning. Its core lies in achieving deep complementarity and intelligent fusion of texture and geometric information through a collaborative mechanism of "two-way guidance".
[0099] First, in the basic feature integration stage, the fusion decision layer receives a 512-dimensional texture feature vector from the 2D feature extraction branch and a 256-dimensional geometric feature vector from the 3D feature extraction branch. These two vectors are directly concatenated along the channel dimension to form a 768-dimensional fused feature vector. This operation breaks down the information barrier of a single modality, constructing a unified high-dimensional feature space that simultaneously contains both "defect appearance texture" and "three-dimensional geometric shape," laying the data foundation for subsequent refined analysis.
[0100] Subsequently, the cross-modal dynamic weighting stage, the most crucial innovative element of the fusion decision layer, is achieved by introducing a cross-modal attention mechanism (CMA). This mechanism does not treat all features equally; instead, it dynamically evaluates and assigns importance weights to different modalities and feature dimensions for the current defect identification task. The specific logic of the cross-modal attention mechanism is as follows: it analyzes the 768-dimensional fusion feature vector, assesses the contribution of surface texture features and geometric structure features to distinguishing various typical defects (e.g., cracks, dents, protrusions, edge damage, scratches), and adjusts their importance weights accordingly.
[0101] Cross-modal attention mechanism (CMA) enables deep collaboration between 2D texture and 3D geometric information. Its core lies in a bidirectional guided fusion strategy, specifically manifested as "texture-guided geometric aggregation" and "geometrically constrained texture response".
[0102] Texture-Guided Geometric Aggregation: The cross-modal attention mechanism first utilizes surface texture features extracted from 2D images to generate a spatial saliency weight map, which quantitatively characterizes the probability that each region of the image contains real defects. Specifically, by analyzing the surface texture features in the 2D image, regions that may indicate the presence of defects are identified, and these regions are assigned a saliency weight. This weight is essentially a numerical representation that describes the probability or likelihood of the existence of an actual physical defect within a specific image region. For example, cracks often appear as thin, continuous abrupt changes in grayscale, while depressions exhibit a contrast pattern of dark centers and bright edges. Because these areas typically show grayscale or color variations different from their surroundings, they are given higher saliency weights. Next, this spatial saliency weight map is used as a priori signal to dynamically adjust the aggregation strategy of geometric features within the local neighborhood of the 3D point cloud. This means that during 3D data analysis, for regions marked as more likely to contain defects according to the saliency map, greater emphasis will be placed on the geometric features within these regions, i.e., their sampling rate and aggregation weight will be increased to more accurately capture subtle geometric distortion features. For example, in regions with high texture saliency (e.g., suspected crack paths or recessed edges), higher sampling and aggregation weights are assigned to the corresponding geometric features, thereby effectively amplifying weak but real geometric distortion signals. Thus, texture-guided geometric aggregation significantly improves the detection capability of subtle deformations (e.g., tiny depressions or gentle protrusions less than 0.3 mm deep) in low-contrast, weak-texture, or reflective interference scenes, enhancing sensitivity to minute but important geometric changes and overcoming the problem of missed detections due to weak geometric signals in traditional methods.
[0103] Geometrically Constrained Texture Response: A cross-modal attention mechanism synchronously analyzes depth variation gradients (e.g., local curvature, abrupt changes in normal, height jumps, etc.) in geometric structural features to generate a depth saliency weight map. This depth saliency weight map is used to dynamically adjust the activation intensity of surface texture features. For example, for regions with real physical deformation (e.g., depth drop regions corresponding to depressions with depth differences exceeding 0.5 mm), the texture response at pixel locations in these regions is enhanced to filter out non-physical depth interference in complex texture backgrounds, ensuring these regions receive higher confidence in classification and decision-making processes. Correspondingly, in background regions that, while possessing complex textures, are geometrically smooth and lack obvious anomalies (e.g., areas with only visual interference and no actual deformation, such as oil stains, water stains, and oxidation spots), even if the texture exhibits an "abnormal" pattern, its texture response is actively suppressed. Thus, the geometrically constrained texture response effectively filters out non-physical visual interference caused by complex surface states, significantly reducing the false alarm rate.
[0104] Through the aforementioned bidirectional guided fusion strategy, the multimodal fusion neural network constructs a collaborative discrimination logic of "texture for identifying suspicious points and geometry for verifying authenticity": texture clues efficiently locate potential defect areas, while geometric information rigorously verifies their physical authenticity. This fusion method not only retains highly discriminative appearance clues but also ensures that all high-confidence prediction results have a solid three-dimensional geometric basis, fundamentally improving the accuracy, robustness, and industrial applicability of the detection system. Compared to traditional methods, the false negative rate of this application is reduced by 60% (e.g., from 15% to 5.2%), significantly improving the accuracy and robustness in identifying complex and subtle defects.
[0105] Ultimately, the feature vectors, after dynamic weighted fusion through the cross-modal attention mechanism, are highly focused on the core information that can accurately distinguish various defects, greatly improving the feature separability between different defect categories and laying a solid foundation for subsequent multi-task joint reasoning.
[0106] Multi-task joint reasoning: The feature vector obtained after the above cross-modal dynamic weighted fusion is fed into the multi-task joint reasoning module of the fusion decision layer, which can output the suspected defect result corresponding to the end face of the steel coil. The suspected defect result represents at least one suspected defect, and for each suspected defect, it includes: defect category (e.g., crack, dent, protrusion, scratch, edge damage, etc.), defect location (i.e., the two-dimensional coordinate positioning information of the defect), three-dimensional depth information corresponding to the defect (i.e., the depth or height variation parameters of the defect), and defect confidence level, etc.
[0107] In some embodiments, the multi-task joint reasoning module may include four parallel output heads, which respectively perform four tasks: defect category prediction, defect location regression, three-dimensional depth information calculation, and defect confidence assessment, thus forming a multi-task collaborative reasoning architecture.
[0108] The defect category prediction branch matches the weighted fused features with the "defect patterns" trained on the model during the training phase. Through learning from 20,000 labeled samples (including 30% defect samples), the model has mastered the unique multimodal feature combinations corresponding to each type of defect. Defect classifications can include, for example, cracks, dents, protrusions, edge damage, and scratches. For instance, the "crack pattern" is characterized by the co-activation of high-weight, elongated continuous texture features and abrupt depth changes along a narrow direction, while the "dent pattern" is characterized by a strong correlation between low-brightness texture in the central region and a local overall depth decrease (e.g., depth difference > 0.5 mm). During inference, this defect category prediction branch calculates the matching probability between the current fused features and various typical defect patterns using fully connected layers and a Softmax function, outputting the probability distribution of each defect type. The defect with the highest probability is the final output defect category.
[0109] Defect Location Regression Branch: Determines the precise coordinates of the defect in physical space. For two-dimensional coordinate positioning information, the model predicts the center position and bounding box of the defect region in the image coordinate system, for example, the upper left corner (x1, y1) and the lower right corner (x2, y2). This is then combined with the pixel-physical size mapping relationship established by Zhang's calibration method (calibration error ≤ 0.3 mm) to convert it into actual planar coordinates (X, Y) on the end face of the steel coil. For three-dimensional spatial positioning, the spatiotemporal registration relationship is further utilized (registration error ≤ 0.5 mm) to extract the depth value (Z-axis) corresponding to the defect region from the three-dimensional point cloud, ultimately outputting complete three-dimensional spatial coordinates (X, Y, Z). This can be directly used to measure the depth of the indentation or the spatial orientation of the crack, providing a spatial reference for subsequent geometric analysis.
[0110] The 3D depth information calculation branch quantifies the degree of geometric deformation of defects. Based on weighted fused features, this branch directly regresses and outputs key geometric parameters that match the defect type. For example, for depression-type defects, it outputs their maximum or average depth; for convex-type defects, it outputs their height relative to the reference plane; and for cracks or scratches, it outputs their projected length, principal orientation angle, and depth variation trend along the orientation. These parameters are not simply copied from the original point cloud data, but are predicted end-to-end through the "defect-geometry" mapping relationship learned by the model, resulting in higher robustness and noise resistance.
[0111] The defect confidence assessment branch comprehensively measures the overall reliability of the detection results. Its calculation logic integrates the Softmax probability of the defect category prediction branch (e.g., reflecting the certainty of category judgment), the bounding box regression quality of the defect location regression branch (e.g., the intersection-union ratio (IOU) of the predicted box and the ideal region), and the physical reasonableness of the 3D depth information predicted by the 3D depth information calculation branch (e.g., whether the depth value is within a reasonable range and whether it aligns with textured salient regions). By weightedly fusing the above multi-dimensional indicators, this defect confidence assessment branch outputs a confidence score in the range of 0 to 1 (e.g., [0, 1]). A defect confidence score of 0 indicates that there is absolutely no belief that a real defect exists in the region, meaning the prediction is highly likely caused by noise, background interference, or non-physical texture changes (such as reflection, oil stains, oxidation spots, etc.), constituting a false detection. A defect confidence score of 1 indicates a high degree of certainty that a real physical defect exists in the region, with its surface texture features highly consistent with 3D geometric deformation and conforming to learned typical defect patterns. This confidence score serves not only for internal threshold judgment but also as a screening basis for subsequent verification stages.
[0112] Through the collaborative work of the above four branches, the suspected defect results output by the multi-task joint reasoning module not only include a comprehensive perception of what the suspected defect is, where it is, how deep it is, and whether it is credible.
[0113] Through step S105, the steel coil end face defect detection model can output the suspected defect results corresponding to the steel coil end face based on the input RGB-D multimodal data.
[0114] In practical applications, inputting RGB-D multimodal data into the steel coil end face defect detection model can output defect detection results in a very short time (single inference time can be ≤50 milliseconds).
[0115] Step S107: Based on the suspected defect results and combined with the three-dimensional point cloud data, perform geometric feature consistency verification on the suspected defects to determine the defect detection results corresponding to the end face of the steel coil.
[0116] The steel coil end-face defect detection model outputs suspected defect results corresponding to the steel coil end-face, which is essentially a probabilistic inference based on statistical regularities and feature matching. Although the model has significantly improved its discrimination ability through cross-modal attention mechanism, it is still possible to misjudge non-physical texture anomalies such as oil stains, oxide spots, and water stains as real defects due to limitations such as training data distribution, lighting interference, or surface reflection.
[0117] In some embodiments, in step S107, a two-level post-processing mechanism is required for the suspected defect results output by the steel coil end face defect detection model: primary screening and geometric feature consistency verification, so as to effectively filter out false defects and ensure that the final detection results have both high accuracy and physical interpretability.
[0118] Regarding the initial screening, all suspected defect results output by the steel coil end-face defect detection model are quickly screened based on their defect confidence level.
[0119] As mentioned earlier, the output of suspected defects includes a defect confidence score, which is a value between [0, 1]. Therefore, a preset confidence threshold (e.g., 0.8) is set, and the following judgment is performed: if the confidence score of a suspected defect is ≥0.8, it is retained as a high-confidence candidate region and enters the next stage of geometric verification; conversely, if the confidence score of a suspected defect is <0.8, it is directly eliminated, no longer considered a valid defect candidate, and will not participate in any subsequent verification or reporting process.
[0120] This primary screening mechanism can efficiently filter out a large number of low-quality predictions (such as false responses caused by image noise, blurred edges, or weak textures), significantly reducing the computational load of subsequent geometric verification, while avoiding low-reliability results from interfering with the final decision.
[0121] Geometric feature consistency verification is performed when the confidence level of the defect in the suspected defect results reaches the confidence level threshold. That is, only for high-confidence candidate regions that have passed the initial screening, a strict physical authenticity verification is performed by combining the original 3D point cloud data.
[0122] In some embodiments, verifying the geometric feature consistency of suspected defects may specifically include the following steps: First, based on the location of the suspected defect, a corresponding local point cloud subset is extracted from the three-dimensional point cloud data.
[0123] Specifically, for each suspected defect output by the steel coil end-face defect detection model, based on the suspected defect's two-dimensional position in the image coordinate system (e.g., center position and bounding box), and combined with a pre-calibrated pixel-point cloud mapping relationship (registration error ≤ 0.5 mm), a three-dimensional point cloud space is precisely back-mapped. This allows the location and extraction of a local point cloud subset centered on the defect. The spatial extent of this local point cloud subset is required to cover the suspected defect. For example, the spatial extent of this local point cloud subset is typically set as a circular area with an appropriate diameter (e.g., 8 mm to 12 mm) centered on the defect center, sufficient to cover the geometry of typical defects such as cracks, dents, or protrusions. Of course, the shape and size of the spatial extent of the local point cloud subset can still be varied. For example, in some examples, the spatial extent of the local point cloud subset can also be set as a rectangular or elliptical area centered on the defect center.
[0124] Next, the random sampling consensus algorithm is used to perform plane fitting on the background points in the local point cloud subset to construct a local reference plane that reflects the local pose state of the current steel coil end face.
[0125] In some embodiments, taking the aforementioned spatial range of a local point cloud subset as an example of a circular region with an appropriate diameter centered on the defect center, an annular background region is delineated around the spatial region formed by this local point cloud subset to construct a local reference surface. The inner diameter of the annular background region is slightly larger than the outer boundary of the spatial region formed by the local point cloud subset. The radial width of the annular background region is configured such that: on the one hand, the radial width is wide enough to cover the local stable surface around the defect (i.e., a representative "normal" surface area), thereby accurately reflecting the current local pose state; if the radial width is too small, it is easily affected by local noise or minor disturbances; on the other hand, the radial width should not be too large to avoid introducing interference factors such as distant macroscopic bending, installation tilt, or adjacent defects, which would lead to distortion of the reference surface. For example, the radial width of the annular background region is typically not less than 3 to 5 times the maximum size of the defect region. For instance, the radial width of the annular background region is greater than or equal to 5 centimeters. This value is determined based on statistical analysis of the end-face morphology of thousands of hot-rolled steel coils. At this scale, the local end-face of the steel coil typically exhibits approximately smooth curved surface characteristics and can effectively avoid interference from adjacent defects or edge effects. Of course, the radial width of the annular background region is not limited to this; in practical applications, it can be dynamically adjusted according to the steel coil diameter, surface roughness, and defect type.
[0126] Subsequently, a robust surface fitting of the annular background region is performed using a random sample consensus algorithm. In the random sample consensus algorithm, by iteratively selecting a subset of points, fitting candidate surface models, and evaluating the number of interior points, a small number of outliers (e.g., minor scratches, oxide spots, or metal spatter) that may be mixed into the annular background region are effectively suppressed, ultimately outputting an optimal local reference surface. It is worth noting that the local reference surface here is not limited to a plane. In practical applications, the reference surface model can be adaptively selected based on the magnitude of the fitting residual in the random sample consensus algorithm. For example, a first-order planar model can be used when the fitting residual is small, while a higher-order surface model (e.g., parabolic surface, hyperboloid) can be used when the fitting residual is large, ensuring that the reference surface accurately reflects the local geometry. This dynamic surface modeling strategy enables the reference surface to accurately reflect the actual geometry of the steel coil end face in this local region. Whether it is a flat area, a slight bulge, or a natural curvature, it can be characterized with high precision, thus providing a reliable reference for subsequent depth quantization.
[0127] Finally, based on the constructed local reference surface, the normal distance of each point in the local point cloud subset relative to the local reference surface is calculated, and key geometric indicators are proposed for defect determination.
[0128] In some embodiments, after constructing a local reference plane, the normal distance from each point in the local point cloud subset to the local reference plane is calculated point by point. Specifically, for concave defects, the normal distance is negative; for convex defects, the normal distance is positive.
[0129] Based on this, two key quantitative indicators were further extracted: Maximum depth deviation: refers to the depth difference corresponding to the point farthest from the local reference plane within a local point cloud subset; that is, the maximum depth deviation. , where D max d represents the maximum depth deviation. i This represents the normal distance from the point to the local reference plane. This index reflects the deformation at the most severe point of the defect and is used to identify local extreme features such as sharp pits, cracks, or towering protrusions.
[0130] Mean depth difference: refers to the average depth deviation of all points in a local point cloud subset relative to the local reference plane, i.e., mean depth difference. ,in, d represents the average depth difference. i This represents the normal distance from the point to the local reference plane, and N represents the number of points in the local point cloud sub-concentration. This index characterizes the overall energy intensity of the defect and is more sensitive to defects with a large area but shallow depth (e.g., indentations, scratches).
[0131] In practical applications, a dual threshold judgment is used to determine the authenticity of defects: The maximum depth deviation is greater than or equal to a preset physical defect threshold, and the average depth difference is greater than or equal to a certain proportion of that threshold, i.e. and Only then is the suspected defect determined to be a real defect. Among them, D... th The preset physical defect threshold can be set according to the material type, size, process standards, or production specifications of the steel coil, for example, 0.3 mm, 0.35 mm, 0.4 mm, 0.45 mm, 0.5 mm, 0.55 mm, 0.6 mm, etc. α is a proportionality coefficient, for example, 0.3, 0.4, etc. This dual threshold judgment design is based on a large amount of production line measurement data: a single threshold is prone to missing large-area shallow defects, while the dual constraint can capture local extrema and ensure that the defect has sufficient spatial extensibility, effectively distinguishing between real deformation and random noise.
[0132] Through the aforementioned collaborative mechanism of "confidence filtering + geometric verification," this application elevates defect determination from the probabilistic output of a deep learning model to deterministic confirmation based on physical laws. All ultimately reported defects possess measurable, reproducible, and traceable three-dimensional geometric evidence, with a depression depth measurement error ≤0.1 mm. More importantly, this process ensures that only areas simultaneously satisfying high model confidence and actual physical deformation are identified as defects, fundamentally eliminating false alarms caused by visual interference such as reflections, oil stains, and oxide scale.
[0133] In addition, the verified defect results can be stored in the defect database, supporting time-series trend analysis of continuous inspection data of the same steel coil at different time points, identifying dynamically evolving defects such as crack propagation and indentation deepening, and providing data support for process optimization and predictive maintenance.
[0134] It is evident that step S107, through the organic combination of primary screening and geometric feature consistency verification, not only significantly improves the reliability of the detection results, but also provides key technical support for achieving high-precision and high-robust industrial online quality inspection.
[0135] After three months of continuous testing and verification on the hot rolling production line of a large steel enterprise, the introduction of this geometric consistency verification process reduced the overall false alarm rate from 12.7% to 4.9%, a decrease of over 60%. Simultaneously, the detection rate of micro-dents with a depth ≥0.3 mm increased to 98.2%. This proposed technical solution can reduce manual inspection positions by more than 80%, reduce subsequent processing losses due to missed defect detection, and improve product quality consistency (increasing the rate of superior products by at least 12%).
[0136] Crucially, the technical solution of this application demonstrates excellent robustness in complex industrial environments: under typical harsh conditions such as light intensity variation of ±50%, steel coil surface reflectivity between 30% and 90%, and environmental vibration acceleration ≤2 mm / s², the stability standard deviation of the identification results of the same standard defect sample after repeated testing is ≤3%, which fully meets the requirements of continuous online operation in industrial sites 24 / 7.
[0137] More importantly, through the aforementioned geometric feature consistency verification mechanism, this application elevates defect determination from the probabilistic output of deep learning models to deterministic confirmation based on physical laws, significantly reducing the false alarm rate, while ensuring that all ultimately reported defects have measurable, reproducible, and traceable physical evidence.
[0138] This application innovatively designs a steel coil end-face defect detection model. This model can extract surface texture features and geometric structure features from RGB-D multimodal data, and then dynamically weight and fuse the surface texture features and geometric structure features using channel splicing and cross-modal attention mechanisms. This achieves deep complementarity and synergy of the two types of heterogeneous information at the feature level, improving the speed of steel coil end-face defect detection. More importantly, it significantly improves the accuracy and robustness of identifying complex and subtle defects. Furthermore, a geometric feature consistency verification mechanism based on 3D point cloud data is introduced. By analyzing the geometric features of suspected defect areas, it determines whether they meet the criteria for judging real physical defects, thereby re-verifying the output results of the steel coil end-face defect detection model. This effectively filters out false defects caused by factors such as surface texture changes and lighting interference, reducing the false detection rate and improving the reliability of the detection results.
[0139] Please see Figure 4 The diagram shown is a flowchart of another embodiment of the steel coil end-face defect detection method of this application.
[0140] like Figure 4 As shown, the steel coil end-face defect detection method in this embodiment includes steps S101, S103, S105, S10, and S109, wherein steps S101 to S107 are the same as those described above. Figure 1 The steps in the embodiments are the same, so they will not be described again. Compared with the previous embodiments, the steel coil end face defect detection method in this embodiment further includes step S109.
[0141] The following is a detailed description of step S109.
[0142] Step S109: The defect detection results are sent to the production line control system, which then generates process control instructions based on the defect detection results.
[0143] In step S109, the defect detection results are sent to the production line control system. Based on the defect detection results, the production line control system generates corresponding process control instructions according to the preset process rules.
[0144] For example, for minor defects, the production line control system can trigger an industrial marking machine (e.g., an inkjet printer) to mark the defect location on the steel coil for reference by downstream processes.
[0145] For example, for serious defects, the production line control system can drive a sorting robot arm to sort the non-conforming steel coils out of the conforming product flow channel and send them to the rework or scrap area.
[0146] For example, if the same type of defect occurs repeatedly, the production line control system can automatically send a production line alarm signal, indicating that the process parameters may need to be adjusted.
[0147] In addition, in some embodiments, the production line control system can also generate a structured defect report based on the received defect detection results, including information such as the steel coil ID, defect category, two-dimensional coordinates and depth information of the defect, and severity level, providing valuable data assets for subsequent quality traceability and process optimization.
[0148] The steel coil end-face defect detection method of this application has full-process automated control capability and can be seamlessly integrated into the PLC system to trigger scanning and detection actions, realizing a closed-loop process that combines detection and production control; while existing technologies mostly rely on static image input or fixed cameras, which have limited detection range and low efficiency.
[0149] The defect detection method for steel coil end face of this application outputs defect detection results with clear defect category, location and physical characteristic information. It can directly interact with the production line control system to realize various production line operations such as automatic marking, sorting and alarm of defective steel coils, forming a closed-loop process that combines detection and production control, thereby improving the automation level and industrial application value of steel coil end face quality inspection.
[0150] Please see Figure 5 The diagram shows a flowchart of one embodiment of the steel coil end-face defect detection model training method of this application. Figure 5 As shown, the training method for the steel coil end-face defect detection model includes the following steps: Step S201: Obtain multiple steel coil end face samples and label the defect truth value information for each steel coil end face sample.
[0151] In some embodiments, it is first necessary to collect a large dataset of representative steel coil end face samples. Each steel coil end face sample may contain synchronously acquired two-dimensional image data and three-dimensional point cloud data, and detailed defect ground truth annotations are performed on each steel coil end face sample.
[0152] Specifically, this application constructs a large-scale dataset containing at least 20,000 (or more, even 50,000 or 100,000) steel coil end face samples, of which defect samples account for 30%, to ensure that the model fully learns the discrimination boundary between normal and abnormal patterns. The ground truth information of the defects includes, but is not limited to: defect category, the corresponding defect region in the two-dimensional image, and the corresponding geometric deformation region and its three-dimensional dimensions (length, depth) in the three-dimensional point cloud.
[0153] The defect categories cover common types found in industrial settings, such as cracks, scratches, dents, indentations, edge damage, burrs, and protrusions.
[0154] The location and bounding box of the defect in the two-dimensional image can be accurately extracted by edge detection or template matching.
[0155] The corresponding geometrically deformed regions in the 3D point cloud are labeled with their spatial location and depth values (e.g., the depression depth is 0.8 mm) based on high-precision scanning data, providing a benchmark for geometric consistency verification.
[0156] In practical applications, these defect truth value annotations can be completed manually or with semi-automatic annotation tools to ensure that each defect record has high-precision spatial positioning and classification labels.
[0157] In addition, to improve the generalization ability of the model, the steel coil end face samples constituting the training sample set should be diverse, covering a variety of material categories (e.g., carbon steel, stainless steel, silicon steel), a variety of diameter specifications (e.g., outer diameter from 800 mm to 2500 mm), and a variety of surface conditions (including oxide scale, oil stains, reflection, rolling marks, etc.) to cover the diverse working conditions in actual production.
[0158] Step S203: Generate a training sample set using multiple steel coil end face samples.
[0159] In step S203, for each steel coil end face sample, the two-dimensional image data and three-dimensional point cloud data in the steel coil end face sample are preprocessed and spatiotemporally registered to form a paired RGB-D sample.
[0160] In practical applications, after obtaining the raw data of the labeled steel coil end face samples, the next step is to perform a series of operations on this raw data to generate a high-quality RGB-D sample set. Specifically, this includes the following steps: First, preprocess the two-dimensional image data and the three-dimensional point cloud data.
[0161] To further enhance the model's robustness and generalization ability in complex industrial environments, targeted data augmentation strategies are introduced in the preprocessing workflow, specifically including: Enhancement is applied to the 2D image data by rotation (e.g., ±15°), scaling (0.8x to 1.2x), and adding Gaussian noise (standard deviation σ = 0.05). These operations aim to simulate image disturbances caused by slight vibrations of steel coils, fluctuations in imaging distance, and sensor electronic noise in actual production lines, exposing the model to diverse imaging conditions during the training phase.
[0162] Enhancements are applied to 3D point cloud data: translation (e.g., ±10 mm), rotation (e.g., ±15°), and scaling (0.8x to 1.2x). These geometric perturbations are used to simulate real-world conditions such as mechanical installation tolerances, thermal expansion deformation, or minute displacements of scanning equipment, thereby significantly improving the model's tolerance to changes in point cloud pose and its geometric stability.
[0163] After data augmentation, further refined preprocessing is performed on the augmented 2D image and 3D point cloud respectively.
[0164] Preprocessing of 2D image data: First, adaptive contrast enhancement is performed on the two-dimensional image data, for example, by using the Adaptive Histogram Equalization (AHE) algorithm or the Contrast Limited AHE (CLAHE) algorithm, to improve the contrast of local areas, thereby making the texture features of small defects (such as fine cracks and shallow scratches) more clearly visible.
[0165] Subsequently, noise suppression processing is performed, such as effectively removing high-frequency random noise in the image through median filtering, while preserving key details such as defect edges to the greatest extent.
[0166] Based on this, further edge detection and image cropping are performed. For example, the Canny edge detection operator is used to accurately identify the outer contour of the steel coil end face, and invalid background areas in the image are cropped accordingly, retaining only the core area containing valid end face information, thereby reducing subsequent computational redundancy and focusing on the area to be inspected.
[0167] Preprocessing of 3D point cloud data: For the synchronously acquired 3D point cloud data, statistical filtering is first performed to remove noise. Based on the local neighborhood point distribution characteristics of each point (e.g., mean and standard deviation), abnormal outliers caused by sensor interference or environmental reflection are automatically identified and removed. Subsequently, a voxel grid downsampling method was used to compress the point cloud density, which significantly reduced the amount of data and improved the efficiency of subsequent processing while maintaining the integrity of the geometric structure.
[0168] Furthermore, surface smoothing is performed on the filtered and downsampled point cloud, and the moving least squares (MLS) method is applied to reconstruct and optimize the point cloud surface to enhance the continuity and smoothness of the macroscopic geometry, providing a more stable geometric basis for depth anomaly detection.
[0169] Then, spatiotemporal registration and multimodal fusion are performed: After completing the independent preprocessing of the two-dimensional image data and the three-dimensional point cloud data, the three-dimensional point cloud data is accurately projected into the corresponding two-dimensional image coordinate system using calibration parameters. Through this registration process, strictly aligned RGB-D multimodal data is generated, in which each effective pixel establishes a one-to-one spatial mapping relationship with one or more point cloud points in three-dimensional space.
[0170] The spatiotemporal registration includes two stages: offline calibration and online registration. During the offline calibration phase, a high-precision calibration board (e.g., a checkerboard pattern) is used in the initial deployment stage to calibrate the industrial line scan camera based on a specific calibration method (e.g., the Zhang Zhengyou calibration method). This process obtains its intrinsic parameter matrix K and lens distortion coefficients (including radial and tangential distortion coefficients). Simultaneously, the extrinsic parameters (i.e., rotation matrix R and translation vector T) between the 3D laser line scan camera and the industrial line scan camera are calibrated. This calibration process establishes the precise geometric relationship between the industrial line scan camera and the 3D laser line scan camera in the world coordinate system.
[0171] During the online registration stage, in each scan, the original two-dimensional image is first subjected to distortion correction, and the nonlinear distortion caused by the lens is eliminated using the distortion coefficients obtained from the calibration above. Subsequently, based on the corrected image coordinate system, each three-dimensional space point (Xw, Yw, Zw) is mapped to the two-dimensional image coordinate system according to the following projection model using the intrinsic parameter matrix K and the extrinsic parameter [R|T]: This projection mapping assigns a corresponding depth value Z to each valid pixel (u, v) in a 2D image, thereby constructing a spatially aligned RGB-D image and generating RGB-D multimodal data that simultaneously contains color and depth information. The RGB channels retain the original 2D texture information, while the depth channel provides precise 3D geometric information. This establishes a one-to-one correspondence between 2D pixels and 3D spatial points, ensuring complete consistency between texture and geometric information in physical space, and achieving depth fusion of multimodal data within a unified coordinate framework.
[0172] By performing the above full-process processing on multiple steel coil end face samples, a large-scale, highly aligned, and strongly generalized RGB-D training sample set is finally generated. This set not only covers diverse working conditions but also has excellent noise robustness and geometric consistency, laying a solid data foundation for the efficient training of subsequent multimodal neural networks.
[0173] Step S205: Input the training sample set into the multimodal fusion neural network for forward propagation and output the corresponding defect prediction results.
[0174] In this application, a multimodal fusion neural network is constructed. The multimodal fusion neural network is configured to extract surface texture features from the two-dimensional image data and geometric structure features from the three-dimensional point cloud data, and to dynamically weight and fuse the surface texture features and the geometric structure features through channel splicing and cross-modal attention mechanisms.
[0175] In practical applications, the multimodal fusion neural network can be composed of three main modules: Two-dimensional feature extraction branch: Deep convolutional neural networks with embedded spatial attention mechanisms (e.g., improved ResNet-50) can extract surface texture features from two-dimensional image data.
[0176] 3D Feature Extraction Branch: Employs a hierarchical feature aggregation network based on PointNet++, which can extract geometric structural features from 3D point cloud data.
[0177] Fusion Decision Layer: Responsible for channel splicing and cross-modal attention mechanism, dynamically weighting and fusing two heterogeneous features (i.e., surface texture features and geometric structure features), and performing tasks such as defect classification, defect location localization, 3D depth information calculation and defect confidence assessment based on the fused features.
[0178] During the training phase, the prepared RGB-D sample set is input into the network in batches. After the forward propagation process, the defect prediction result for each sample is obtained. The suspected defect result represents at least one suspected defect and includes the defect category, defect location, three-dimensional depth information corresponding to the defect, and defect confidence for each suspected defect.
[0179] Step S207: Calculate the loss between the defect prediction result and the defect truth information.
[0180] To quantify the difference between the model output and the true label, three types of loss functions are defined: defect classification loss, defect localization regression loss, and geometric consistency loss.
[0181] Defect classification loss: used to measure the cross-entropy error between the predicted class and the actual class; Defect localization regression loss: used to evaluate the distance deviation between the predicted defect location and the actual location; Geometric consistency loss: used to constrain the degree of overlap between the two-dimensional defect region and the three-dimensional depth anomaly region predicted by the model in spatial projection, ensuring that the two have consistency.
[0182] These three types of losses are combined to form an overall objective function, which guides the direction of model parameter updates.
[0183] Step S209: Based on the loss, the backpropagation algorithm is used to update the network parameters of the multimodal fusion neural network, and the optimization is iteratively performed until the model converges, thus obtaining the trained steel coil end face defect detection model.
[0184] Gradient descent (e.g., the Adam optimizer) is used to adjust the network weights based on the calculated total loss. Through iterative training, the objective function value is gradually reduced, enabling the model to accurately capture the multimodal features of different types of defects, eventually reaching convergence and forming a high-performance steel coil end-face defect detection model.
[0185] In a real-world deployment case at a large steel company, the model built using the training method proposed in this invention performed exceptionally well during three consecutive months of online testing. Compared to traditional visual inspection solutions, this model not only significantly improved the recognition rate of complex and subtle defects (such as microcracks less than 0.2 mm in width) but also drastically reduced the false alarm rate, achieving an overall detection accuracy of over 98%, thus providing strong technical support for automated quality control on the production line.
[0186] In summary, the steel coil end-face defect detection model training method disclosed in this application, through carefully designed data preprocessing, multimodal feature extraction, dynamic weighted fusion, and loss function optimization strategies, has successfully achieved efficient and accurate identification of various defects on the steel coil end face, and has significant industrial application value.
[0187] This application also discloses a steel coil end-face defect detection system. Please refer to [link / reference]. Figure 6 The image shown is a schematic diagram of the steel coil end-face defect detection system of this application in one embodiment.
[0188] like Figure 6 As shown, the steel coil end face defect detection system of this application includes: a data acquisition module 41, a data processing module 43, a data analysis module 45, and a verification module 47.
[0189] The data acquisition module 41 includes a motion actuator and a two-dimensional image acquisition device and a three-dimensional point cloud acquisition device mounted on the motion actuator. The motion actuator drives the two-dimensional image acquisition device and the three-dimensional point cloud acquisition device to dynamically scan the end face of the steel coil to be tested along a preset trajectory, and simultaneously acquire two-dimensional image data and three-dimensional point cloud data.
[0190] In some embodiments, the motion actuator may be, for example, a six-axis industrial robot, the two-dimensional image acquisition device may be, for example, a high-resolution industrial line scan camera, and the three-dimensional point cloud acquisition device may be, for example, a high-precision 3D laser line scan camera. By controlling the six-axis industrial robot, the industrial line scan camera and the 3D laser line scan camera can be driven to dynamically scan the end face of the steel coil to be tested along a preset trajectory, and simultaneously acquire two-dimensional image data and three-dimensional point cloud data.
[0191] The data processing module 43 is used to preprocess and spatiotemporally register the two-dimensional image data and the three-dimensional point cloud data to generate paired RGB-D multimodal data, wherein the pixels in the two-dimensional image and the points in the three-dimensional point cloud have a spatial correspondence.
[0192] The data analysis module 45 is equipped with a steel coil end face defect detection model. The model receives the RGB-D multimodal data and extracts surface texture features and geometric structure features from the RGB-D multimodal data. The surface texture features and geometric structure features are dynamically weighted and fused through channel splicing and cross-modal attention mechanisms. Based on the weighted and fused features, the suspected defect results of the steel coil end face are output. The suspected defect results include defect category, defect location, and defect confidence.
[0193] The verification module 47 is used to perform geometric feature consistency verification on the suspected defect based on the suspected defect result and the three-dimensional point cloud data, and to determine the defect detection result corresponding to the end face of the steel coil.
[0194] In some embodiments, geometric feature consistency verification of the suspected defects is performed when the defect confidence in the suspected defect results reaches a confidence threshold. That is, only for high-confidence candidate regions, a strict physical authenticity verification is performed by combining the original 3D point cloud data.
[0195] It should be noted that the steel coil end-face defect detection system disclosed in the above embodiments belongs to the same concept as the aforementioned steel coil end-face defect detection method. The specific operation methods of each module and unit have been described in detail in the method embodiments and will not be repeated here. In practical applications, the steel coil end-face defect detection system disclosed in the above embodiments can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above. This is not a limitation here.
[0196] The steel coil end-face defect detection system of this application works collaboratively through four modules: data acquisition, processing, analysis and verification. It can adapt to steel coils of different sizes and deformations, supports fully automatic online detection, has a low false detection rate and strong robustness, significantly improves detection accuracy and reliability, and greatly enhances the level of intelligent quality control in steel coil production.
[0197] This application also discloses an electronic device, such as... Figure 7 The diagram shown illustrates the structure of an electronic device according to an embodiment of this application. This electronic device can be a terminal or a server, specifically: The electronic device includes a processor 501 and a memory 503. The processor 501 and the memory 503 can communicate via a bus 502. The memory 503 can store program instructions. The processor 501 implements the steps in the steel coil end face defect detection method in the previous embodiment by running the program instructions in the memory 503.
[0198] Bus 502 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, although only one thick line is used in the diagram, this does not indicate that there is only one bus or one type of bus.
[0199] The processor 501 can be implemented as a central processing unit (CPU), microprocessor unit (MCU), system on chip (System on Chip), or field programmable logic array (FPGA).
[0200] The memory 503 may include volatile memory for temporary data storage during operation, such as random access memory (RAM).
[0201] The memory 503 may also include non-volatile memory for data storage, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state disk (SSD).
[0202] In some embodiments, the electronic device may further include a communication interface 504. The communication interface 504 is used for communication with external devices. In specific examples, the communication interface 504 may include one or more wired and / or wireless communication circuit modules. For example, the communication interface 504 may include one or more of, such as a wired network card, a USB module, a serial interface module, etc. The wireless communication protocols followed by the wireless communication module include, for example, Nearfield communication (NFC) technology, Infrared (IR) technology, Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), Bluetooth (BT), Global Navigation Satellite System (GNSS), etc., one or more of these.
[0203] Electronic devices may also include input units and output units. The input unit can be used to receive input digital or character information, and to generate keyboard, mouse, touchscreen, joystick, optical, or trackball signal inputs related to user settings and function control. The output unit can be used to present or transmit processing results, status information, and interactive feedback to the user, for example, through a display screen (such as an LCD screen or touchscreen), indicator lights, speakers, vibration modules, printers, or other audible, visual, or tactile output devices, generating text, image, sound, or tactile signals corresponding to equipment operating status, recipe dispensing results, operation prompts, or alarm information.
[0204] Specifically, in this embodiment, the processor 501 in the electronic device loads the executable files corresponding to the processes of one or more computer programs into the memory 503 according to the following instructions, and the processor 501 runs the program instructions stored in the memory 503 to realize the various steps of the aforementioned steel coil end face defect detection method.
[0205] For details on the implementation of each of the above steps, please refer to the previous examples, which will not be repeated here.
[0206] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by a computer program, or by a computer program controlling related hardware. The computer program can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0207] Therefore, this application further discloses a computer-readable storage medium storing a computer program that can be loaded by a processor to execute the steps in any of the steel coil end-face defect detection methods disclosed in this application.
[0208] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0209] Since the computer program stored in the computer-readable storage medium can execute the steps in any of the steel coil end-face defect detection methods disclosed in the embodiments of this application, the beneficial effects that any of the steel coil end-face defect detection methods disclosed in the embodiments of this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.
[0210] This application also discloses a computer program product, which includes a computer program stored in a computer-readable storage medium. A processor of an electronic device reads the computer program from the computer-readable storage medium and executes the computer program, causing the electronic device to perform the methods disclosed in various optional implementations of the above-described steel coil end-face defect detection method.
[0211] The foregoing has provided a detailed description of a method for detecting end-face defects in steel coils, a system for detecting end-face defects in steel coils, an electronic device, and a computer-readable storage medium disclosed in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and its core ideas and are not intended to limit this application. Any person skilled in the art can modify or change the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or changes made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.
Claims
1. A method for detecting defects on the end face of steel coils, characterized in that, Includes the following steps: The control motion actuator drives the two-dimensional image acquisition device and the three-dimensional point cloud acquisition device to dynamically scan the end face of the steel coil to be tested along a preset trajectory, and simultaneously acquire two-dimensional image data and three-dimensional point cloud data. The two-dimensional image data and the three-dimensional point cloud data are preprocessed and spatiotemporally registered to generate paired RGB-D multimodal data, wherein the pixels in the two-dimensional image and the points in the three-dimensional point cloud have a spatial correspondence. The RGB-D multimodal data is input into a pre-trained steel coil end-face defect detection model, and the following operations are performed: surface texture features and geometric structure features are extracted from the RGB-D multimodal data; the surface texture features and geometric structure features are dynamically weighted and fused using channel splicing and cross-modal attention mechanisms; based on the weighted fused features, the suspected defect results of the steel coil end face are output, wherein the suspected defect results represent at least one suspected defect and include the defect category, defect location, corresponding 3D depth information, and defect confidence level for each suspected defect; and Based on the suspected defect results, and in conjunction with the three-dimensional point cloud data, the geometric feature consistency of the suspected defects is verified to determine the defect detection result corresponding to the end face of the steel coil.
2. The method for detecting defects on the end face of steel coils according to claim 1, characterized in that, The control motion actuator drives the two-dimensional image acquisition device and the three-dimensional point cloud acquisition device to dynamically scan the end face of the steel coil to be tested along a preset trajectory, simultaneously acquiring two-dimensional image data and three-dimensional point cloud data, including: The steel coil to be tested is transported to the testing area, and the testing process is triggered by the position detection component; The control motion actuator drives the configured two-dimensional image acquisition device and three-dimensional point cloud acquisition device to scan the end face of the steel coil along a preset trajectory, wherein the angle between the optical axis of the two-dimensional image acquisition device and the optical axis of the three-dimensional point cloud acquisition device and the normal of the end face of the steel coil is less than or equal to 30°. Through hardware synchronization signals, the two-dimensional image acquisition device and the three-dimensional point cloud acquisition device can synchronously acquire two-dimensional image data and three-dimensional point cloud data during the scanning process, with a synchronization accuracy of less than or equal to 1 microsecond; In the operation of controlling the motion actuator to drive the two-dimensional image acquisition device and the three-dimensional point cloud acquisition device to dynamically scan the end face of the steel coil to be measured along a preset trajectory: The preset trajectory is a spiral path or a grid path, and the scanning speed is 0.5 m / s to 2 m / s; The resolution of the two-dimensional image data is greater than or equal to 4096×4096 pixels; The point density of the three-dimensional point cloud data is greater than or equal to 100 points / square centimeter; Specifically, when the diameter of the steel coil is greater than or equal to a preset diameter threshold and the flatness of the end face meets the preset conditions, the preset trajectory is planned using a spiral scanning path; when the diameter of the steel coil is less than the preset diameter threshold and the flatness of the end face does not meet the preset conditions, the preset trajectory is planned using a grid scanning path.
3. The method for detecting defects on the end face of steel coils according to claim 1, characterized in that, Preprocessing the two-dimensional image data includes: performing adaptive contrast enhancement processing on the two-dimensional image data, using a filtering algorithm to remove noise, and extracting the outline of the steel coil end face through edge detection to crop invalid areas; The preprocessing of the three-dimensional point cloud data includes: removing outliers by statistical filtering, reducing the point cloud density by downsampling, and then smoothing the downsampled point cloud. Spatiotemporal registration of the two-dimensional image data and the three-dimensional point cloud data includes: projecting the three-dimensional point cloud data onto the two-dimensional image coordinate system using calibration parameters to generate RGB-D multimodal data with depth information; The dynamic weighted fusion of the surface texture features and the geometric structure features through channel splicing and cross-modal attention mechanisms includes: The surface texture features and the geometric structure features are concatenated in the channel dimension; By utilizing the spatial saliency weights from the surface texture features, the aggregation weight allocation from the geometric structure features is guided to enhance geometric distortion features in low-contrast regions; and By utilizing the depth variation weights of the geometric structural features, the texture response of suspected regions in the surface texture features is enhanced to filter out non-physical depth interference in complex texture backgrounds.
4. The method for detecting defects on the end face of steel coils according to claim 1, characterized in that, The steel coil end-face defect detection model is formed based on a multimodal fusion neural network trained on a multimodal fusion neural network, which includes: The two-dimensional feature extraction branch employs a deep convolutional neural network with embedded spatial attention mechanism to extract surface texture features from the two-dimensional image data. The 3D feature extraction branch employs a hierarchical feature aggregation point cloud processing network to extract geometric structure features from the 3D point cloud data; and The fusion decision layer is used to stitch the surface texture features and the geometric structure features in the channel dimension, and to perform dynamic weighted fusion based on the stitched features through a cross-modal attention mechanism. Then, based on the weighted fused features, defect category classification, defect location regression, three-dimensional depth information calculation and defect confidence assessment are performed simultaneously to generate suspected defect results corresponding to the end face of the steel coil.
5. The method for detecting defects on the end face of a steel coil according to claim 1, characterized in that, Geometric feature consistency verification of the suspected defects is performed when the defect confidence level in the suspected defect results reaches a confidence threshold, specifically including: Based on the location of the suspected defect, extract the corresponding local point cloud subset from the 3D point cloud data; The random sampling consensus algorithm is used to perform plane fitting on the background points in the local point cloud subset to construct a local reference plane that reflects the local pose state of the current steel coil end face; Calculate the normal distance of each point in the local point cloud subset relative to the local reference plane, and extract the maximum depth value and the average depth difference; and If the maximum depth value or the average depth difference is greater than the preset physical defect threshold, the suspected defect is determined to be a real defect.
6. The method for detecting defects on the end face of a steel coil according to claim 1, characterized in that, After determining the defect detection result corresponding to the end face of the steel coil, the method further includes: sending the defect detection result to the production line control system, and the production line control system generating process control instructions based on the defect detection result to execute any one or more of the following operations: automatic marking, sorting, and production line alarm.
7. A training method for a steel coil end-face defect detection model, characterized in that, Includes the following steps: Multiple steel coil end face samples are acquired. Each steel coil end face sample contains synchronously acquired two-dimensional image data and three-dimensional point cloud data. Defect truth value information is labeled for each steel coil end face sample. The defect truth value information includes defect category, corresponding defect area in two-dimensional image, and corresponding geometric deformation area in three-dimensional point cloud. A training sample set is generated using multiple steel coil end face samples; wherein, for each steel coil end face sample, the two-dimensional image data and the three-dimensional point cloud data in the steel coil end face sample are preprocessed and spatiotemporally registered to form paired RGB-D samples. The training sample set is input into a multimodal fusion neural network for forward propagation, and the corresponding defect prediction result is output. The multimodal fusion neural network is configured to extract surface texture features from the two-dimensional image data and geometric structure features from the three-dimensional point cloud data, and to dynamically weight and fuse the surface texture features and the geometric structure features through channel splicing and cross-modal attention mechanisms. Calculate the loss between the defect prediction result and the defect ground truth information; the loss includes defect classification loss, defect localization regression loss, and geometric consistency loss, wherein the geometric consistency loss is used to constrain the overlap between the two-dimensional defect region and the three-dimensional depth anomaly region predicted by the model in spatial projection; and Based on the aforementioned loss, the backpropagation algorithm is used to update the network parameters of the multimodal fusion neural network, and the optimization is iteratively performed until the model converges, thus obtaining a trained steel coil end face defect detection model.
8. The training method for the steel coil end-face defect detection model according to claim 7, characterized in that, It also includes the following steps: A multimodal fusion neural network is constructed, comprising a two-dimensional feature extraction branch, a three-dimensional feature extraction branch, and a fusion decision layer. The two-dimensional feature extraction branch employs a deep convolutional neural network with embedded spatial attention mechanism to extract surface texture features from the two-dimensional image data. The three-dimensional feature extraction branch employs a point cloud processing network with hierarchical feature aggregation to extract geometric structure features from the three-dimensional point cloud data. The fusion decision layer concatenates the surface texture features and the geometric structure features along the channel dimension, and dynamically weights and fuses the concatenated features using a cross-modal attention mechanism.
9. A steel coil end face defect detection system, characterized in that, include: The data acquisition module includes a motion actuator and a two-dimensional image acquisition device and a three-dimensional point cloud acquisition device mounted on the motion actuator. The motion actuator drives the two-dimensional image acquisition device and the three-dimensional point cloud acquisition device to dynamically scan the end face of the steel coil to be tested along a preset trajectory, and simultaneously acquire two-dimensional image data and three-dimensional point cloud data. The data processing module is used to preprocess and spatiotemporally register the two-dimensional image data and the three-dimensional point cloud data to generate paired RGB-D multimodal data, wherein the pixels in the two-dimensional image and the points in the three-dimensional point cloud have a spatial correspondence. The data analysis module is configured with a steel coil end-face defect detection model. This model receives RGB-D multimodal data and extracts surface texture and geometric features from it. The surface texture and geometric features are dynamically weighted and fused using channel splicing and cross-modal attention mechanisms. Based on the weighted fused features, the module outputs suspected defect results for the steel coil end-face. Each suspected defect result represents at least one suspected defect and includes the defect category, defect location, corresponding 3D depth information, and defect confidence level for each suspected defect. The verification module is used to perform geometric feature consistency verification on the suspected defect based on the suspected defect results and the three-dimensional point cloud data, and to determine the defect detection result corresponding to the end face of the steel coil.
10. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, it implements the steel coil end face defect detection method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Steel coil end face defect detection method based on deep learning
CN109829900A
Steel coil end face defect detection method
CN115294039A
Steel coil end face quality inspection system based on machine vision technology
CN115901782A
Steel coil end face scanning device and method based on 3D imaging technology
CN117871535A
Detection method and system for steel coil single-side overflow defect and storage medium
CN118279225A
Cited By
Multi-mode collaborative awareness and time-space fusion equipment coil steel tallying system
CN122067063A
A waterlogging detection method based on horizontal plane features, medium and equipment
CN122199547A
A waterlogging detection method based on horizontal plane features, medium and equipment
CN122199547B