A steel plate surface defect detection method and system based on multi-source data fusion
By using water stain pretreatment and multi-source data fusion technology, combined with data acquired by 3D and 2D cameras, and utilizing neural networks to extract features and generate structured reports, the problems of water stain interference and multi-source data collaborative utilization in steel plate surface defect detection are solved, achieving high-precision and high-robust automated detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ANHUI YANSHI INTELLIGENT TECHNOLOGY CO LTD
- Filing Date
- 2025-12-08
- Publication Date
- 2026-05-29
Smart Images

Figure CN121482497B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of automated inspection technology in steel production, and in particular to a method and system for detecting surface defects in steel plates based on multi-source data fusion. Background Technology
[0002] In steel production, steel plates are prone to various defects such as cracks, holes, pits, and scale indentation during rolling, cooling, and transportation. Their surface quality directly affects the subsequent processing performance and safety of the product. Currently, mainstream inspection methods mainly rely on manual visual inspection or machine vision technology based on a single 2D image.
[0003] With technological advancements, 3D point cloud technology has provided a new approach for extracting three-dimensional surface features. However, existing technologies suffer from the following shortcomings in the data processing layer: First, water stain interference is a significant issue. Residual cooling water from the production process can severely impact data acquisition, generating abnormal noise in 3D point cloud data and creating reflective artifacts in 2D images. Existing technologies lack effective water stain preprocessing mechanisms, potentially leading to increased false positive and false negative rates in humid environments. Second, the ability to identify three-dimensional defects is insufficient. Traditional 2D image detection methods cannot accurately acquire surface depth information, and the detection of three-dimensional defects such as pits and protrusions may have inherent limitations. The existing methods suffer from several limitations, including a high false negative rate. Furthermore, the collaborative utilization of multi-source data is insufficient, and current methods fail to effectively combine the complementary advantages of 3D point clouds and 2D images: 3D point clouds excel at geometric morphology representation but are weak in texture recognition, while 2D images are adept at texture capture but lack depth information. The lack of an effective cross-modal feature fusion mechanism may lead to insufficient stability in detecting complex defects such as water-stained cracks. In addition, the adaptive capability of data preprocessing is lacking; existing systems mostly use fixed parameter configurations, making it difficult to adapt to the inspection needs of steel plates with different specifications and surface conditions. Especially in high-speed production line environments, it is impossible to achieve an effective balance between processing efficiency and detection accuracy. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method and system for detecting surface defects of steel plates based on multi-source data fusion. By eliminating water stain interference, unifying multi-source data benchmarks, optimizing data quality, accurately identifying defects, and realizing automated feedback and data closure, a high-precision and robust steel plate surface defect detection system covering the entire detection process is constructed.
[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:
[0006] Firstly, a method for detecting surface defects in steel plates based on multi-source data fusion, the method comprising:
[0007] The steel plate surface is pretreated with water stains to reduce the residual humidity to below a predetermined low humidity threshold, thereby obtaining a dry surface.
[0008] Based on the dry surface, multi-source data are collected by synchronously triggered 3D laser line scan camera and 2D industrial line scan camera to obtain line scan laser point cloud data, line scan reflectivity grayscale image, line scan depth grayscale image and area array reflectivity color image, establish spatial correspondence between multi-source data, and form multi-source dataset based on the spatial correspondence between multi-source data.
[0009] Based on a multi-source dataset, a multi-source data calibration region is established on the surface of a steel plate and partitioned into a grid to obtain a grid framework. The multi-source data corresponding to the grid framework is analyzed to obtain preprocessing parameter adjustment values for image enhancement and point cloud filtering. Using the preprocessing parameter adjustment values, adaptive bilateral filtering and grayscale stretching are performed on the line scan reflectivity grayscale image, the line scan depth grayscale image, and the area array reflectivity color image. At the same time, outlier removal is performed on the line scan laser point cloud data to obtain a high-quality preprocessed multi-source dataset.
[0010] The pre-processed high-quality multi-source dataset is input into a pre-trained multi-data fusion neural network to extract texture and geometric features, which are then fused through a cross-modal attention mechanism to obtain the detection results.
[0011] Based on the detection results, a structured defect report is obtained and transmitted to the PLC system via the OPC UA protocol to trigger automatic marking and quality grading operations. At the same time, the detection data is stored in a blockchain database for quality traceability and model iteration.
[0012] Secondly, a steel plate surface defect detection system based on multi-source data fusion includes:
[0013] The surface pretreatment module is used to pretreat the water stains on the steel plate surface, reduce the residual humidity to below a predetermined low humidity threshold, and obtain a dry surface.
[0014] The acquisition module is used to acquire multi-source data based on the dry surface through a synchronously triggered 3D laser line scan camera and a 2D industrial line scan camera, to obtain line scan laser point cloud data, line scan reflectivity grayscale image, line scan depth grayscale image and area array reflectivity color image, establish spatial correspondence between multi-source data, and form a multi-source dataset based on the spatial correspondence between multi-source data.
[0015] The processing module is used to establish a multi-source data calibration region on the surface of a steel plate based on a multi-source dataset and perform partitioning and meshing to obtain a mesh framework; analyze the multi-source data corresponding to the mesh framework to obtain preprocessing parameter adjustment values for image enhancement and point cloud filtering; use the preprocessing parameter adjustment values to perform adaptive bilateral filtering and grayscale stretching on the line scan reflectivity grayscale image, the line scan depth grayscale image, and the area array reflectivity color image, and simultaneously perform outlier removal on the line scan laser point cloud data to obtain a high-quality preprocessed multi-source dataset;
[0016] The defect analysis module is used to input the pre-processed high-quality multi-source dataset into a pre-trained multi-data fusion neural network, extract texture and geometric features, and obtain the detection results after fusion through a cross-modal attention mechanism;
[0017] The results feedback module is used to generate a structured defect report based on the detection results, which is transmitted to the PLC system via the OPC UA protocol to trigger automatic marking and quality grading operations. At the same time, the detection data is stored in the blockchain database for quality traceability and model iteration.
[0018] Thirdly, a computing device, comprising:
[0019] One or more processors;
[0020] A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method.
[0021] Fourthly, a computer-readable storage medium storing a program that, when executed by a processor, implements the method.
[0022] The above-described solution of the present invention has at least the following beneficial effects:
[0023] By controlling residual humidity to below a predetermined low humidity threshold through water stain pretreatment, interference factors are eliminated at the data acquisition source, avoiding abnormal noise in point clouds and image reflection artifacts caused by water stains. This provides a clean and stable surface foundation for subsequent multi-source data acquisition, ensuring the authenticity and purity of the original data. Simultaneous triggering of two types of cameras to acquire multiple types of data ensures temporal consistency of data while taking into account geometric morphology and texture details. Establishing spatial correspondences between multi-source data and forming a dataset achieves a unified spatial benchmark for different modal data, providing structurally sound and clearly correlated data support for cross-modal fusion. Partitioned gridding decomposes the data into refined local units, focusing on subtle data differences and avoiding global analysis from obscuring features. Adaptive generation of preprocessing parameters allows filtering and enhancement operations to adapt to different regions. Data features: Simultaneous optimization of image and point cloud quality, preserving defect details while suppressing noise, resulting in high-quality data that balances integrity and clarity; parallel extraction of texture and geometric features via dual branches, fully leveraging the core value of multi-source data; cross-modal attention mechanism to strengthen intermodal correlation features and weaken irrelevant interference, enabling fused features to comprehensively characterize defect attributes; output of complete detection information including category, location, and size, providing data basis for subsequent operations; structured defect reports adapted to the transmission and parsing needs of industrial systems, ensuring smooth data flow; triggering automated marking and grading, improving the automation level of production processes; blockchain storage to establish tamper-proof quality traceability records, integrating data to provide samples that fit real-world scenarios for model iteration, forming a data closed loop to enhance comprehensive utilization value. Attached Figure Description
[0024] Figure 1 This is a schematic flowchart of a steel plate surface defect detection method based on multi-source data fusion provided by an embodiment of the present invention.
[0025] Figure 2 This is a schematic diagram of a steel plate surface defect detection system based on multi-source data fusion, provided by an embodiment of the present invention. Detailed Implementation
[0026] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0027] like Figure 1 As shown in the figure, an embodiment of the present invention proposes a method for detecting surface defects in steel plates based on multi-source data fusion. The method includes the following steps:
[0028] Step 100: Perform water stain pretreatment on the surface of the steel plate to reduce the residual humidity to below the predetermined low humidity threshold and obtain a dry surface;
[0029] Step 200: Based on the dry surface, multi-source data is collected by synchronously triggered 3D laser line scan camera and 2D industrial line scan camera to obtain line scan laser point cloud data, line scan reflectivity grayscale image, line scan depth grayscale image and area array reflectivity color image, establish spatial correspondence between multi-source data, and form multi-source dataset based on spatial correspondence between multi-source data.
[0030] Step 300: Based on the multi-source dataset, a multi-source data calibration region is established on the steel plate surface and partitioned into a grid to obtain a grid framework; the multi-source data corresponding to the grid framework is analyzed to obtain preprocessing parameter adjustment values for image enhancement and point cloud filtering; using the preprocessing parameter adjustment values, adaptive bilateral filtering and grayscale stretching are performed on the line scan reflectivity grayscale image, the line scan depth grayscale image, and the area array reflectivity color image, while outlier removal is performed on the line scan laser point cloud data to obtain a high-quality preprocessed multi-source dataset;
[0031] Step 400: Input the preprocessed high-quality multi-source dataset into the pre-trained multi-data fusion neural network to extract texture and geometric features, and then fuse them through a cross-modal attention mechanism to obtain the detection results;
[0032] Step 500: Based on the detection results, a structured defect report is obtained and transmitted to the PLC system via the OPC UA protocol to trigger automatic marking and quality grading operations. At the same time, the detection data is stored in the blockchain database for quality traceability and model iteration.
[0033] In this embodiment of the invention, water stains on the steel plate surface are removed to reduce residual humidity and minimize interference from water stains on subsequent data acquisition, laying the foundation for high-quality data acquisition. Simultaneous acquisition of multiple types of data establishes spatial correspondences between data, integrates three-dimensional geometric information and two-dimensional texture information, enriches data dimensions, and provides complete and closely related data support for multi-source data fusion. Preprocessing parameters are optimized through partitioned grid division, adaptively improving image detail clarity, effectively filtering outliers in point clouds, enhancing the purity and consistency of multi-source data, and strengthening the identifiability of intrinsic data features. Texture and geometric features are fully extracted, and deep feature fusion is achieved through a cross-modal attention mechanism, strengthening the correlation between different types of defect features and improving the comprehensive extraction effect of defect information. Structured defect information is generated, enabling efficient linkage between detection results and the production system, ensuring the security and traceability of data storage, providing data support for continuous model optimization, and improving the practicality and sustainability of the detection system.
[0034] In a preferred embodiment of the present invention, step 100, which involves pretreating the steel plate surface to reduce residual humidity to below a predetermined low humidity threshold to obtain a dry surface, includes:
[0035] Step 101: The surface humidity of the steel plate entering the detection area is monitored by an infrared humidity sensor to obtain a surface humidity status signal. Specifically, this includes: first, calibrating the monitoring range of the infrared humidity sensor to match the width of the channel through which the steel plate enters the detection area, ensuring that the sensor can fully cover the area to be detected on the steel plate surface; then, setting the real-time data acquisition frequency of the sensor to match the transmission speed of the steel plate production line, so as to achieve continuous capture of humidity data at different locations on the steel plate surface; during operation, the sensor will convert the physical quantity of humidity on the steel plate surface sensed by the sensor through an internal signal conversion unit, converting the analog signal into a digital signal that can be recognized by the subsequent processing unit, thereby forming a complete surface humidity status signal, providing basic data support for subsequent liquid water interference determination.
[0036] Step 102: Process the surface humidity status signal. When liquid water interference is detected, a water removal device activation signal is obtained. Specifically, this includes: first, setting a threshold for liquid water interference humidity detection. This threshold needs to be determined in conjunction with the anti-interference requirements of subsequent 3D point cloud acquisition and 2D image acquisition. Basic data is obtained through multiple sets of comparative experiments. Specifically, 3D point cloud data and 2D image data of the steel plate surface are collected under different humidity conditions. The noise rate of the 3D point cloud and the proportion of reflection artifacts in the 2D image corresponding to different humidity levels are statistically analyzed. When the humidity drops to a certain value, the noise rate of the 3D point cloud can be stably less than or equal to 3%, and... The initial critical value is set at 10% or less for 2D image reflection artifacts. Considering the differences in physical properties of steel plate surfaces under different production scenarios such as cold rolling and hot rolling, the above-mentioned comparative experiments need to be conducted for different steel plate specifications to obtain critical value ranges suitable for different scenarios. For different steel plate specifications, such as thickness of 6 to 20 mm and width of 500 to 2000 mm, the lower limit of this range is set as the general threshold for judging humidity interference from liquid water. If a specific production line is to be selected, the critical value of the corresponding scenario of the production line is selected as the judgment threshold to ensure that the threshold setting has both universality and scenario adaptability.
[0037] Subsequently, the digital data in the surface humidity status signal obtained in step 101 is analyzed in real time. First, according to the monitoring area division rules of the steel plate surface, the monitoring range of the infrared humidity sensor is divided into several monitoring zones matching the width of the steel plate. For example, when the steel plate width is 1800mm, it is divided into 6 monitoring zones of 300mm each. 3 to 5 evenly distributed monitoring points are set in each monitoring zone. During the analysis process, the humidity digital data of each monitoring point needs to be extracted one by one according to the preset analysis frequency. The preset analysis frequency is adapted to the transmission speed of the steel plate production line. For example, when the production line speed is 0.3m / s, the analysis frequency is set to be greater than or equal to 5Hz. At the same time, the validity of the extracted humidity data is verified, and abnormal data caused by instantaneous fluctuations of the sensor is eliminated. Specifically, when the difference between two consecutive humidity data collected at a certain point exceeds ±10% of the previous data, the data is determined to be an abnormal value and the average of the two valid data is used to replace it, ensuring that the analyzed data can truly reflect the humidity of the steel plate surface.
[0038] Next, the humidity values at each point after validity verification are compared with the preset liquid water interference judgment threshold one by one in real time. The comparison process follows the rule of combining zone judgment and overall judgment: for each monitoring zone, if the humidity values of two or more consecutive monitoring points in the zone are higher than the judgment threshold, the zone is judged to have liquid water interference; when two or more monitoring zones on the entire steel plate surface are judged to have liquid water interference, or when more than half of the monitoring points in a single zone have humidity values higher than the judgment threshold, the entire steel plate surface is judged to have liquid water interference.
[0039] When the determination result indicates the presence of liquid water interference, the signal generation module is activated. Within 100ms of receiving the determination result, this module generates a water removal device start signal and simultaneously embeds the location information of the monitoring point exceeding the threshold (including the monitoring zone to which it belongs and its specific coordinates within the zone) into the start signal. This allows the subsequent water removal device to adjust its operating parameters accordingly based on the location information, providing location-guided instructions for the accurate triggering and efficient execution of subsequent water removal operations.
[0040] Step 103: Based on the activation signal of the dehydration device, control the high-pressure air blowing unit and the mechanical scraping unit to perform coordinated dehydration operations to obtain the steel plate surface after preliminary dehydration. Specifically, this includes: after receiving the activation signal of the dehydration device generated in step 102, firstly, acquiring the actual width specifications and surface humidity distribution data of the steel plate to be treated, and simultaneously determining the vertical distance from the nozzle outlet of the high-pressure air blowing unit to the steel plate surface, using this distance as the height of the cone; then, based on the cone lateral area algorithm, determining the lateral range that the airflow needs to cover on the steel plate surface with the steel plate width as a reference, and corresponding this lateral coverage range to the arc length parameter of the cone base, calculating the effective coverage area of the high-pressure airflow to ensure that the effective coverage area can completely cover the area to be dehydrated on the steel plate surface, avoiding local water stains due to incomplete airflow coverage; based on this, combining the location and humidity value of the high humidity area in the surface humidity distribution data, setting the output air pressure value of the high-pressure air blowing unit according to the matching relationship between the effective coverage area and the airflow density, so that the high humidity area can obtain sufficient airflow intensity within the effective coverage area to quickly remove water stains, while the low humidity area is treated with an appropriate airflow. To avoid excessive air blowing and resource waste, the blowing time of the high-pressure air blowing unit is set based on the airflow coverage length parameter calculated from the lateral area of the cone and the transmission speed of the steel plate production line. This ensures that each position on the surface of the steel plate can remain within the airflow coverage area for a sufficient time during transmission, achieving full effect on water stains in different locations. Simultaneously, the moving speed of the scraper blade of the mechanical scraping unit is set according to the transmission speed of the steel plate production line, and the moving path of the scraper blade is synchronized with the effective coverage area of the high-pressure airflow. The effective coverage area is the area corresponding to the lateral area of the cone, ensuring that the squeegee maintains a suitable contact force with the steel plate surface, guaranteeing the squeegee effect while avoiding damage to the steel plate surface. Then, operation commands are sent to the high-pressure air blowing unit and the mechanical scraping unit respectively, controlling the high-pressure air blowing unit to output high-pressure airflow according to the set air pressure value and duration, and the mechanical scraping unit to perform scraping operations along the steel plate surface according to the set moving speed. The two units work together to complete the preset operation cycle, achieving efficient removal of liquid water from the steel plate surface, and finally obtaining a preliminarily dehydrated steel plate surface.
[0041] Step 104: Real-time humidity monitoring is performed on the surface of the steel plate after preliminary dehydration. When the monitoring data indicates that the residual humidity has reached a predetermined low humidity threshold, a work termination signal is obtained and the obtained dry surface is confirmed. Specifically, this includes: First, determining the predetermined low humidity threshold. The setting of this threshold is based on ensuring that there is no interference between subsequent 3D point cloud and 2D image acquisition. This is accomplished through multiple sets of controlled experiments and scenario adaptation verification. Steel plate samples matching the actual application scenario are selected, covering different surface treatment types such as cold rolling, hot rolling, and galvanizing, and common specifications with thicknesses of 6 to 20 mm and widths of 500 to 2000 mm. The testing conditions of the production line are simulated in a laboratory environment, and the surface humidity of the steel plate is controlled at different gradient values such as 1%, 3%, 5%, 7%, and 9%. For each humidity gradient, 3D point cloud data and 2D image data are simultaneously acquired, statistically analyzed, and classified. The noise rate of 3D point clouds and the proportion of reflection artifacts in 2D images were analyzed under different humidity gradients. When the humidity gradient decreased to 5%, the noise rate of 3D point clouds was stably less than or equal to 3%, and the proportion of reflection artifacts in 2D images was less than or equal to 10%. Under this condition, the false detection rate of subsequent defect detection could be controlled within 5%. This 5% was used as the basic low humidity threshold for general scenarios. At the same time, supplementary experiments were conducted for different production line speeds and steel plate surface conditions. For different production line speeds, such as 0.3m / s, 1m / s, and 2m / s, and steel plate surface conditions such as oxide scale and coating, if the production line speed was greater than or equal to 2m / s or the steel plate surface was galvanized, the experiment verified that the humidity needed to be further controlled to less than or equal to 4.5% to ensure data acquisition without interference. At this time, the basic threshold was dynamically fine-tuned for this specific scenario, and finally the predetermined low humidity threshold adapted to the current detection scenario was determined.
[0042] Subsequently, the infrared humidity sensor is activated to continuously collect humidity data on the surface of the steel plate after preliminary dehydration. The collection frequency is adapted to the transmission speed of the steel plate production line. For example, when the production line speed is 0.3m / s, the collection frequency is set to greater than or equal to 8Hz to ensure that the residual humidity changes at different locations on the steel plate surface can be captured in real time. The collected residual humidity data is processed by the signal conversion module inside the sensor, converting the analog humidity sensing signal into a 16-bit precision digital signal, forming a digital signal that can be recognized by the data processing unit. The digital signal is analyzed in real time according to the preset analysis rules. That is, abnormal data caused by instantaneous sensor jitter is first eliminated. When the humidity data of a certain collection point differs from the mean of the five adjacent collection points by more than ±0.5%, it is judged as an abnormal value and replaced by the mean of the adjacent points. Then, the average humidity value of each monitoring zone is calculated. Each monitoring zone is consistent with the monitoring zone divided in step 102, and this is used as the actual residual humidity value of the zone.
[0043] The actual residual humidity value of each monitoring zone is continuously compared with the predetermined low humidity threshold. The comparison process must meet the dual requirements of temporal continuity and spatial full coverage. Temporally, it must be within 10 consecutive collection cycles, where each collection cycle is the reciprocal of the collection frequency. The actual residual humidity value of all monitoring zones must be less than or equal to the predetermined low humidity threshold. Spatially, it must be ensured that all monitoring zones meet the threshold requirement, and no zone's actual residual humidity value exceeds the threshold. When both of the above requirements are met, it is determined that the actual residual humidity value has stably met the standard.
[0044] At this time, the trigger signal generation module runs. Within 50ms after receiving the stable compliance judgment result, the module generates a water removal device operation termination signal. At the same time, it transmits the final residual humidity value and compliance time of each monitoring zone to the system control unit through the data feedback channel. The control unit confirms that the steel plate surface has met the drying requirements, providing a stable and interference-free surface environment for the subsequent accurate acquisition of multi-source data from the 3D laser line scanning camera and the 2D industrial line scanning camera.
[0045] In a preferred embodiment of the present invention, step 200 above involves acquiring multi-source data based on the dried surface using a synchronously triggered 3D laser line scan camera and a 2D industrial line scan camera. This yields line scan laser point cloud data, line scan reflectivity grayscale images, line scan depth grayscale images, and area array reflectivity color images. Spatial correspondences between the multi-source data are established, and a multi-source dataset is formed based on these spatial correspondences, including:
[0046] Step 201: Based on the dry surface, the 3D laser line scanning camera and the 2D industrial line scanning camera are controlled by a synchronous trigger signal to synchronously acquire raw 3D sensing data and 2D image data. Specifically, this includes: First, determining the core configuration of the synchronous trigger signal based on the real-time operating parameters of the steel plate production line. The real-time transmission speed of the steel plate is obtained through the production line PLC system. This speed range is suitable for the high-speed production line requirements mentioned in the background technology, typically 0.3 m / s to 2 m / s. The synchronous trigger frequency is calculated based on this transmission speed. For example, when the production line speed is 0.3 m / s, the trigger frequency is set to be greater than or equal to 5 Hz to ensure that both the 3D laser line scanning camera and the 2D industrial line scanning camera can complete full-area data acquisition without omissions during the process of each steel plate passing through the detection area. Second, the acquisition parameters of the two types of cameras are calibrated collaboratively. The 3D laser line scanning camera uses a 532nm laser emitter, and its scanning... The scanning frequency was set to 300Hz to match the trigger signal, and the laser power was adjusted to 60% to avoid excessive laser reflection from the dry surface, which could lead to data distortion. The 2D industrial line scan camera used a 12K resolution CCD sensor, with the scanning frequency set to greater than or equal to 20kHz and the exposure time adjusted to 80μs to ensure consistency with the acquisition timing of the 3D camera. Subsequently, the synchronous trigger module was activated. This module generates a synchronous pulse signal based on the preset trigger frequency, with the signal transmission delay controlled to less than or equal to 1μs. The signal is sent to the control units of the 3D laser line scan camera and the 2D industrial line scan camera, respectively. After receiving the signal, the two types of cameras synchronously start the acquisition action, capturing the three-dimensional geometric information and two-dimensional texture information of the dry steel plate surface in real time. Finally, the initial raw 3D sensing data and 2D image data are obtained. The raw 3D sensing data includes laser distance signal and reflection intensity signal, and the 2D image data includes RGB color signal.
[0047] Step 202 involves reconstructing the point cloud and analyzing the reflectivity of the original 3D sensing data to obtain line-scan laser point cloud data, line-scan reflectivity grayscale images, and line-scan depth grayscale images. Specifically, this includes: firstly, reconstructing the laser distance signal in the original 3D sensing data by calling the preset intrinsic parameters (including focal length, pixel size, and distortion coefficient) and extrinsic parameters (including the 45° angle between the camera optical axis and the steel plate surface, and the installation height parameter) of the 3D laser line-scan camera. This converts the distance signal of each laser scanning point into three-dimensional spatial coordinate values, where the X-axis corresponds to the production line transmission direction, the Y-axis corresponds to the steel plate width direction, and the Z-axis corresponds to the vertical direction of the steel plate surface. Based on this, the point density of the coordinate data is optimized by supplementing sparse areas using a neighborhood interpolation algorithm, ensuring that the final generated line-scan laser point cloud data has a point density greater than or equal to 200 points / cm². 2First, the reflection intensity signal in the original 3D sensing data is analyzed to meet the accuracy requirements of subsequent geometric feature extraction. Second, the reflection intensity value, which is usually a digital quantity from 0 to 4095, is converted into a grayscale value from 0 to 255 through grayscale mapping rules. During the mapping process, outlier values are truncated. In particular, outlier values, such as values below 100 or above 3800, are then filtered by mean to eliminate small fluctuations, generating a line scan reflectivity grayscale image with a resolution of 1000×3200. Finally, the Z-axis coordinate value of the line scan laser point cloud data is analyzed for depth. The effective range of the Z-axis coordinate is also mapped to a grayscale value from 0 to 255. The effective range of the Z-axis coordinate is usually from -5mm to 5mm. With the steel plate reference surface as the zero point, the smaller the Z-axis coordinate (concave), the lower the grayscale value, and the larger the Z-axis coordinate (convex), the higher the grayscale value. After edge smoothing, a line scan depth grayscale image with a resolution of 1000×3200 is generated.
[0048] Step 203 involves color space conversion and pixel calibration of the 2D image data to obtain a surface reflectance color image. Specifically, this includes: first, color space conversion of the original 2D image data. The original data is typically stored in RGB color space. To eliminate the impact of uneven illumination on texture recognition on the dry surface, the RGB color signal is converted to HSV color space. During the conversion, the hue (H) channel value remains unchanged, while the saturation (S) channel value is increased by 10% overall to enhance texture contrast. The lightness (V) channel value is adaptively adjusted according to the illumination intensity of different areas on the steel plate surface. For example, the V value is reduced by 5% to 8% in areas with strong illumination and increased by 8% to 12% in areas with weak illumination to ensure uniform overall color brightness. Next, pixel calibration is performed. A high-precision checkerboard calibration plate (10mm×10mm grid size) is used to pre-acquire distortion parameters of the 2D industrial line scan camera, including radial and tangential distortion coefficients. Based on these distortion parameters, the coordinates of each pixel in the 2D image data are corrected, adjusting the position of the distorted pixels to their actual corresponding positions on the steel plate surface, with the correction error controlled to be less than or equal to 1 pixel. Subsequently, reflectivity feature enhancement is performed on the calibrated image. By calculating the ratio of the HSV value of each pixel to the HSV value of the standard white pixel, the relative reflectivity of the pixel is obtained. The relative reflectivity is then mapped back to the RGB color space to generate a color image with a resolution greater than or equal to 12000×8000, ensuring that the texture features of the steel plate surface, such as oxide scale and scratches, are clearly distinguishable in the image.
[0049] Step 204: Based on the line-scanned laser point cloud data and the area array reflectivity color image, establish the spatial correspondence between the line-scanned laser point cloud data and the area array reflectivity color image through a pre-calibrated sensor pose transformation matrix. Specifically, this includes: First, the sensor pose transformation matrix needs to be pre-calibrated. This calibration process requires standard instruments and standardized operations. The specific steps are as follows: First, select a standard calibration block with three-dimensional coordinate markers. The calibration block needs to meet the requirement that the number of feature points is greater than or equal to 20, and the three-dimensional coordinates of each feature point are known quantities. The coordinate accuracy needs to be controlled within ±0.01mm. At the same time, each feature point has a high-contrast mark on its surface that facilitates image recognition, such as a black circular mark with a diameter of 2mm, to ensure accurate positioning during subsequent image acquisition. Then, place the standard calibration block at the reference position of the steel plate in the detection area. This reference position needs to be aligned with the center of the transmission path of the steel plate during actual detection, and the upper surface of the calibration block needs to be flush with the reference surface of the steel plate during detection to avoid distortion of the calibration results due to placement deviation.
[0050] After placing the calibration block, the 3D laser line scan camera is activated to acquire laser point cloud data of the standard calibration block. During the acquisition process, the parameters of the 3D camera are kept consistent with those used in subsequent actual testing, such as laser power at 60% and scanning frequency at 300Hz. After acquisition, based on the geometric features of the point cloud data, such as the spherical protrusion structure and corner shape of the feature points, the three-dimensional coordinates of each known feature point in the 3D point cloud coordinate system are identified and extracted through point cloud segmentation and feature extraction algorithms. During the extraction process, the coordinates of each feature point need to be measured three times, and the average value is taken as the final three-dimensional coordinate value of the feature point to reduce measurement error.
[0051] Simultaneously, a 2D industrial line scan camera is activated to acquire a color image of the reflectance of the standard calibration block. The acquisition parameters are kept consistent with those used in actual testing, such as an exposure time of 80μs and a scanning frequency of ≥20kHz. After acquisition, an image feature matching algorithm, such as a SIFT-based feature point matching algorithm, is used to identify the high-contrast markers of each feature point in the image. This determines the two-dimensional pixel coordinates of each feature point in the 2D image pixel coordinate system. The two-dimensional coordinates of each feature point are also repeatedly identified and extracted three times, and the average value is taken as the final two-dimensional pixel coordinate value.
[0052] Based on the extracted 3D coordinates and 2D pixel coordinates of the feature points, the pose transformation matrix is calculated using the principle of spatial coordinate transformation (such as the Perspective-n-Point algorithm). This matrix must include translation and rotation parameters. Specifically, the translation parameters are the offsets of the 3D point cloud coordinate system relative to the 2D image pixel coordinate system in the X, Y, and Z directions, and the rotation parameters are the rotation angles of the 3D point cloud coordinate system relative to the 2D image pixel coordinate system around the X, Y, and Z axes. To further improve the matrix accuracy, the calibration process requires at least 5 repeated acquisitions and calculations. The average value of the pose transformation matrix parameters obtained from each calculation is taken. The final matrix must meet the requirement that the transformation error is less than or equal to 0.1 mm, that is, the deviation between the coordinates of any feature point after transformation by this matrix and the actual coordinates must be less than or equal to 0.1 mm, ensuring that the matrix can accurately realize the transformation between the two coordinate systems.
[0053] After the sensor pose transformation matrix is pre-calibrated, it can be used to establish the spatial correspondence between the linear laser point cloud data and the area array reflectivity color image: for each three-dimensional coordinate point in the linear laser point cloud data, its coordinate value is substituted into the pre-calibrated pose transformation matrix for calculation. Through the translation and rotation parameters in the matrix, the two-dimensional pixel coordinates of the three-dimensional coordinate point in the area array reflectivity color image are obtained; at the same time, for each pixel in the area array reflectivity color image, its pixel coordinates are substituted into the pose transformation matrix in reverse for inverse operation, and the three-dimensional spatial coordinate range of the pixel in the 3D point cloud coordinate system is calculated, that is, the three-dimensional coordinate interval corresponding to the steel plate surface area covered by the pixel.
[0054] To verify the accuracy of the spatial correspondence, at least 100 feature points were randomly selected for verification. Specifically, 100 feature points that were not involved in the matrix calculation were selected, or 100 feature points were randomly selected from the selected feature points. Their corresponding coordinates were obtained through matrix transformation and compared with the actual coordinates to ensure that the correspondence accuracy between the three-dimensional coordinates and the two-dimensional pixel coordinates was greater than or equal to 99.5%. This ensured the establishment of a stable spatial mapping relationship between the line scan laser point cloud data and the area array reflectivity color image.
[0055] Step 205: Based on the spatial correspondence, and utilizing the inherent index association between the line-scanned reflectivity grayscale image, the line-scanned depth grayscale image, and the line-scanned laser point cloud data, all multi-source data are unified to the same spatial coordinate system, forming a multi-source dataset with a unified spatial correspondence. Specifically, this includes: first, determining a unified spatial coordinate system, with the center point of the steel plate entry end of the detection station as the origin, the production line transmission direction as the positive X-axis, the steel plate width direction as the positive Y-axis, and the vertical upward direction from the steel plate surface as the positive Z-axis. The unit of length in the coordinate system is set to mm, with precision retained to one decimal place; secondly, based on the data established in step 204... Spatial correspondence is established by transforming the two-dimensional pixel coordinates of the area array reflectivity color image into a three-dimensional coordinate range in a unified coordinate system using a pose transformation matrix. This ensures that each pixel in the image corresponds to a specific spatial region within the unified coordinate system. Simultaneously, the inherent index association between the line-scanned reflectivity grayscale image, the line-scanned depth grayscale image, and the line-scanned laser point cloud data is utilized. Since these three types of data are synchronously acquired by the same 3D laser line-scanning camera, they have a one-to-one correspondence between frame and pixel indices. Each pixel in the line-scanned reflectivity grayscale image and the line-scanned depth grayscale image is associated with the corresponding three-dimensional coordinate point index in the line-scanned laser point cloud data through index matching, thus mapping it to the unified coordinate system. Subsequently, spatial alignment verification of the multi-source data is performed. Natural feature points on the steel plate surface, such as edge inflection points and minor surface pits, are selected. The coordinates or pixel information of these feature points are extracted from the four types of data: line-scanned laser point cloud data, line-scanned reflectivity grayscale image, line-scanned depth grayscale image, and area array reflectivity color image. Their consistency in the unified coordinate system is verified. If the deviation exceeds 0.1 mm, it is corrected by fine-tuning the index association parameters. Finally, the four types of data are integrated and sorted according to the X-axis and Y-axis coordinates of a unified coordinate system. Each spatial coordinate point corresponds to the storage of the three-dimensional coordinate value, line scan reflectivity gray value, line scan depth gray value, and area array reflectivity color value of the line scan laser point cloud, forming a well-structured multi-source dataset, which provides a unified data benchmark for cross-modal feature extraction of subsequent deep learning models.
[0056] In a preferred embodiment of the present invention, step 300 involves establishing a multi-source data calibration region on the steel plate surface based on a multi-source dataset and dividing it into partitioned grids to obtain a grid framework; analyzing the multi-source data corresponding to the grid framework to obtain preprocessing parameter adjustment values for image enhancement and point cloud filtering; and using the preprocessing parameter adjustment values, performing adaptive bilateral filtering and grayscale stretching on the line-scanned reflectivity grayscale image, the line-scanned depth grayscale image, and the area array reflectivity color image, while simultaneously removing outliers from the line-scanned laser point cloud data to obtain a preprocessed high-quality multi-source dataset, including:
[0057] Step 301: Based on the unified spatial coordinate system of the multi-source dataset, extract the boundary coordinates of the effective detection area on the steel plate surface; calculate the geometric center position and normal vector direction of the detection area based on the boundary coordinates of the effective detection area on the steel plate surface. Specifically, this includes: First, relying on the unified spatial coordinate system established by the multi-source dataset, extract the boundary coordinates of the effective detection area on the steel plate surface; specifically, first, select points in the line scan laser point cloud data whose Z-axis coordinates are within a reasonable range of steel plate thickness from the multi-source dataset. For example, when the steel plate thickness is 6 to 20 mm, the Z-axis coordinate is limited to within ±10 mm of the reference plane. At the same time, combine the areas in the area array reflectance color image where the pixel gray value is within the texture feature range of the steel plate surface, exclude the black background area and the bright interference area, and spatially superimpose the areas selected by the two types of data. Take the superimposed area on the X-axis (production line transmission direction) and Y-axis (steel plate transmission direction). The extreme values along the width of the plate are used as the boundary coordinates of the effective detection area. That is, the X-axis boundary is the minimum and maximum X-coordinates of all points in the superimposed area, and the Y-axis boundary is the minimum and maximum Y-coordinates of all points in the superimposed area. After obtaining the boundary coordinates, the geometric center position of the detection area is calculated. Specifically, the arithmetic mean of the minimum and maximum values of the X-axis boundary is used as the center X-coordinate, the arithmetic mean of the minimum and maximum values of the Y-axis boundary is used as the center Y-coordinate, and the Z-axis coordinate is taken as the Z-value of the steel plate reference surface, which together constitute the geometric center position. At the same time, the normal vector direction is calculated. By selecting more than 20 feature points evenly distributed in the effective detection area, such as edge inflection points and sampling points in the surface flat area, the spatial coordinate matrix of these points is constructed, the covariance matrix of the matrix is solved, and its eigenvector is calculated. The eigenvector that is consistent with the direction perpendicular to the steel plate surface is the normal vector direction of the detection area.
[0058] Step 302: Establish a local coordinate system for the steel plate surface based on the geometric center position and normal vector direction; calculate the boundary coordinates of the scanning coverage area in the local coordinate system according to the horizontal and vertical scanning field of view of the 3D laser line scanning camera. Specifically, this includes: establishing a local coordinate system for the steel plate surface based on the calculation of the geometric center position and normal vector direction; taking the geometric center position as the origin of the local coordinate system, setting the normal vector direction as the positive Z-axis of the local coordinate system, and setting the production line transmission direction, which is consistent with the X-axis of the unified spatial coordinate system, as the positive X-axis of the local coordinate system; determining that the positive Y-axis of the local coordinate system is consistent with the width direction of the steel plate using the right-hand rule, ensuring that the local coordinate system fits the shape of the steel plate surface; then calculating the boundary of the scanning coverage area according to the horizontal and vertical scanning field of view of the 3D laser line scanning camera. Coordinates: First, obtain the camera's hardware parameters, including the horizontal scanning field of view α, which typically ranges from 30° to 60°, the vertical scanning field of view β, which typically ranges from 5° to 15°, and the vertical distance d from the camera lens center to the geometric center of the steel plate surface, which is determined according to the installation parameters and is usually 500 to 800 mm. In the local coordinate system, calculate the half-width of the horizontal (Y-axis) scan coverage as d × tan(α / 2) and the half-length of the vertical (X-axis) scan coverage as d × tan(β / 2). With the origin of the local coordinate system as the center, extend the horizontal half-width in the positive and negative Y-axis directions and the vertical half-length in the positive and negative X-axis directions respectively to obtain the X-axis and Y-axis boundary coordinates of the scan coverage area in the local coordinate system. The Z-axis boundary coordinate is consistent with the steel plate reference surface.
[0059] Step 303: Determine the spatial range of the multi-source data calibration area on the steel plate surface based on the boundary coordinates of the scan coverage area, and establish the multi-source data calibration area. Specifically, this includes: determining the spatial range of the multi-source data calibration area on the steel plate surface based on the calculation results of the aforementioned scan coverage area boundary coordinates; specifically, taking the intersection area of the effective detection area boundary obtained in step 301 and the scan coverage area boundary obtained in step 302; taking the overlapping interval of the effective detection area X-axis boundary and the scan coverage area X-axis boundary in the X-axis direction, that is, the X-axis calibration boundary is the larger of the two X-axis minimum values and the smaller of the two X-axis maximum values; similarly, in the Y-axis direction, taking the larger of the two Y-axis minimum values and the smaller of the two Y-axis maximum values; the Z-axis direction is still within ±5mm of the steel plate reference surface, ensuring that the calibration area covers the range of effective data that the camera can actually collect, and focuses on the core detection area of the steel plate surface, avoiding interference from invalid background data or unscanned areas to subsequent processing, and finally forming the complete spatial range parameters of the multi-source data calibration area.
[0060] Step 304: Based on the spatial range of the multi-source data calibration area, calculate the reference grid cell size that meets the requirements of point cloud data feature analysis. Specifically, this includes: calculating the reference grid cell size that meets the requirements of point cloud data feature analysis based on the spatial range of the multi-source data calibration area; firstly, determining the basic requirements for point cloud data feature analysis, i.e., a single grid cell must contain a sufficient number of point cloud data (usually greater than or equal to 10 points) to support the calculation of density distribution features, while avoiding excessively large cells that might obscure local subtle features; firstly, calculating the average point density of the line-scan laser point cloud data in the multi-source data within the calibration area, such as a known point density greater than or equal to 200 points / cm². 2 The minimum area of a single grid cell is calculated based on the average point density, i.e., minimum area = number of points required / average point density. For example, if the required number of points is 12 and the average point density is 200 points / cm², then... 2 Therefore, the minimum area = 12 / 200 = 0.06 cm² 2 Combining the X and Y axis lengths of the calibration area, the minimum area is converted into the side length of an approximate square. If the minimum area is 0.06 cm²... 2 The side length is approximately 0.245cm, which is equivalent to 2.45mm. To facilitate subsequent calculations and division, it is rounded down to 2.5mm. The final reference grid unit size is determined to be 2.5mm × 2.5mm, ensuring that the number of point clouds in a single unit can meet the requirements of feature analysis and reflect the differences in point cloud distribution in local areas.
[0061] Step 305: Based on the reference mesh cell size, calculate the number of mesh divisions along the longitudinal and transverse directions of the steel plate surface, respectively. Specifically, this includes: calculating the number of mesh divisions along the longitudinal (X-axis direction) and transverse (Y-axis direction) directions of the steel plate surface based on the determined reference mesh cell size; first, obtain the X-axis length L of the multi-source data calibration area. x With Y-axis length L y The value is calculated from the boundary coordinates determined in step 303, that is, the maximum value of the X-axis minus the minimum value of the X-axis, which is L. x The difference between the maximum value on the Y-axis and the minimum value on the Y-axis is L. y Number of vertical divisions N x The calculation method is L x Divide by the baseline mesh cell side length. If the calculation result has a decimal, round it up. At the same time, fine-tune the cell side length (the fine-tuning amount is less than or equal to 0.1 mm) to ensure N x It is an integer, such as L x When the unit length is 2000mm and the unit side length is 2.5mm, N x =2000 / 2.5=800, no fine-tuning needed; number of horizontal divisions N y The calculation method is consistent with that of the vertical direction, i.e., L yDivide by the baseline mesh cell side length, round up, and fine-tune the cell side length to ensure N y It is an integer, such as L y When the unit length is 1800mm and the unit side length is 2.5mm, N y =1800 / 2.5=720, finally obtaining the vertical N x With horizontal N y The number of grid divisions is determined to ensure that the grid completely covers the calibration area after division.
[0062] Step 306: Based on the number of grid divisions, establish uniformly distributed regular geometric units within the multi-source data calibration region to obtain a grid framework that perfectly matches the spatial range of the multi-source data calibration region. Specifically, this includes: based on the vertical N... x With horizontal N y The number of grid divisions is determined, and uniformly distributed regular geometric units are established within the multi-source data calibration area; firstly, starting from the minimum X-axis value (X...) of the calibration area... min ) and the minimum value of the Y-axis (Y min Starting from the positive X-axis, divide the area into Nx units, with each unit having an X-axis range of X. min +(i-1)×cell side length to X min +i×cell side length, where i is from 1 to N x Integers; divide N sequentially along the positive Y-axis. y There are 1 unit, and the Y-axis range of each unit is Y. min +(j-1)×unit side length to Y min +j×cell side length, where j is from 1 to N y The integer; the spatial range of each regular geometric unit is determined by the corresponding X-axis interval and Y-axis interval, while the Z-axis interval remains ±5mm from the steel plate reference surface, forming N. x ×N y The grid consists of uniformly distributed rectangular grid cells. After the grid is divided, the maximum X-axis value of all grid cells is verified to be consistent with the maximum X-axis value of the calibration area, and the maximum Y-axis value is consistent with the maximum Y-axis value of the calibration area. This ensures that the grid frame is fully matched with the spatial range of the calibration area, with no areas missing or exceeding the limit.
[0063] Step 307: Based on the aforementioned mesh framework, calculate the density distribution characteristics of the point cloud data and the gray-level gradient characteristics of the image data within each mesh cell. Specifically, this includes: based on the established mesh framework, calculating the density distribution characteristics of the point cloud data and the gray-level gradient characteristics of the image data within each mesh cell; for the point cloud data density distribution characteristics, traversing each mesh cell, counting the total number of points of the line-scanned laser point cloud data falling within the spatial range of that cell, dividing the total number of points by the area of that cell (i.e., multiplying the cell side length by the cell side length) to obtain the point cloud density value of that cell, and using the point cloud density values of all cells as the density of the point cloud data. Gray-level gradient characteristics: For the gray-level gradient characteristics of image data, the processing is carried out separately for line scan reflectance gray-level images, line scan depth gray-level images, and area array reflectance color images (after being converted to gray-level images). Taking the image region corresponding to a single grid cell as the object, the gray-level difference between each pixel in the region and its horizontal and vertical adjacent pixels is calculated. The square root of the sum of the squares of the horizontal and vertical gray-level differences of each pixel is used to obtain the gradient value of that pixel. All pixels in the cell are traversed and the distribution of gradient values is statistically analyzed, such as the maximum value, minimum value, and median of the gradient values, which are used as the gray-level gradient characteristics of the image data of that cell.
[0064] Step 308: Normalize the density distribution characteristics of the point cloud data to obtain normalized density coefficients; perform amplitude statistics on the gray-level gradient characteristics of the image data to obtain the average gradient intensity. Specifically, this includes: normalizing the point cloud density values of all grid cells to obtain normalized density coefficients; firstly, finding the maximum value ρ among all cell point cloud density values. max With minimum value ρ min For the density value ρ of each cell i According to (ρ) i -ρ min ) / (ρ max -ρ min The normalized density coefficient of the unit is calculated in a way that ensures the coefficient value is between 0 and 1, eliminating the influence of the difference in absolute density values between different units and making the density features have a unified comparison standard. At the same time, the amplitude of the gray-level gradient features of the image data of each unit is statistically analyzed to obtain the average gradient intensity. For the image gradient value set of each unit, the arithmetic mean of all gradient values is calculated. If there are abnormal pixels with a gradient value of 0, such as background areas, these abnormal values are removed first and then the average value is calculated to obtain the average gradient intensity of the unit.
[0065] Step 309 involves weighting and fusing the normalized density coefficient and average gradient intensity according to a predetermined weight ratio to obtain a weighted fusion result. Specifically, this includes weighting and fusing the normalized density coefficient and average gradient intensity of each grid cell according to a predetermined weight ratio. The determination of this predetermined weight ratio must be based on the influence of multi-source data features on the subsequent preprocessing effect and must be calibrated through multiple sets of control experiments to ensure balance. Specifically, steel plate samples covering different defect types such as pits, cracks, and oxide scale indentation, as well as different surface humidity states, are selected. Multiple weight combinations are designed for control experiments, with the weight ratios of the normalized density coefficient and average gradient intensity in each weight combination being 0.3:0.7, 0.4:0.6, and 0.5:0.5, respectively. For each weight combination, subsequent image enhancement and point cloud filtering preprocessing are performed on the multi-source data of the samples. Then, the point cloud denoising effect and image enhancement effect of the preprocessed data are detected. The point cloud denoising effect includes outlier removal rate and effective defect point retention rate, while the image enhancement effect includes texture detail clarity and noise suppression degree. The comprehensive processing effect corresponding to each weight is statistically analyzed. After multiple rounds of experimental verification, it was found that when the weight of the normalized density coefficient is 0.4 and the weight of the average gradient intensity is 0.6, it can ensure that abnormal noise is accurately removed during the point cloud filtering process without losing effective geometric features, and ensure that the defect texture details are clearly preserved during the image enhancement process without excessively amplifying noise. The preprocessing effect of the two types of data reaches the optimal balance state. Therefore, this weight ratio is determined as the predetermined weight ratio.
[0066] After determining the predetermined weight ratio, for each grid cell, the weighted fusion value is calculated by multiplying the normalized density coefficient by 0.4 and adding the average gradient strength by 0.6. After the fusion values of all cells are calculated, these values need to be range-checked to determine whether each fusion value is within a reasonable range of 0 to 1. If there is a fusion value less than 0, it is truncated to 0; if there is a fusion value greater than 1, it is truncated to 1. Through this verification and adjustment process, the validity and consistency of the fusion results of all cells are ensured, providing a balanced and reliable feature basis for the subsequent synchronous generation of preprocessing parameter adjustment values for image enhancement and point cloud processing.
[0067] Step 310: Based on the weighted fusion result, simultaneously generate preprocessing parameter adjustment values for image enhancement and point cloud processing. Specifically, this includes: based on the weighted fusion result of each grid cell, simultaneously generating preprocessing parameter adjustment values for both image enhancement and point cloud processing; firstly, establishing a mapping relationship between the fusion result and the parameter adjustment values. This mapping relationship is determined through previous experimental calibration. When the fusion result approaches 1, it indicates high point cloud density, large image gradient, and rich details. The parameter adjustment value for image enhancement needs to approach 0.8, corresponding to a smaller filtering intensity to preserve details. The parameter adjustment value for point cloud processing needs to approach 0.9, corresponding to... A smaller outlier threshold is used for accurate noise reduction. When the fusion result approaches 0, it indicates low point cloud density, small image gradient, and blurred details. The parameter adjustment value for image enhancement needs to approach 0.3, corresponding to a larger filtering intensity to suppress noise. The parameter adjustment value for point cloud processing needs to approach 0.2, corresponding to a larger outlier threshold to avoid erroneous deletion of valid points. For the fusion result of each unit, the image enhancement parameter adjustment value and point cloud processing parameter adjustment value corresponding to that unit are calculated by linear interpolation according to the above mapping relationship. This achieves synchronous generation of the two types of parameters, ensuring that the parameter adjustment is accurately matched with the multi-source data features within the unit, and avoiding insufficient adaptability caused by independent parameter generation.
[0068] Step 311: Based on the preprocessing parameter adjustment values of the image enhancement processing, calculate the spatial domain sigma value and gray-level domain sigma value corresponding to the line scan reflectance grayscale image, the line scan depth grayscale image, and the area array reflectance color image, respectively; dynamically configure the bilateral filter parameters for each image using the spatial domain sigma value and the gray-level domain sigma value; perform filtering processing on each image using the configured bilateral filter parameters to obtain noise-suppressed image data, specifically including: adjusting the image enhancement processing parameters based on each grid cell, Calculate the spatial domain sigma value and gray-level domain sigma value for each of the three image types. First, establish a correlation rule between parameter adjustment values and sigma values: the larger the parameter adjustment value, the smaller the spatial domain sigma value. For example, an adjustment value of 0.8 corresponds to a spatial sigma of 0.6, and an adjustment value of 0.3 corresponds to a spatial sigma of 1.5. The gray-level sigma value decreases as the adjustment value increases; for example, an adjustment value of 0.8 corresponds to a gray-level sigma of 0.5, and an adjustment value of 0.3 corresponds to a gray-level sigma of 1.2. Based on this rule, for each... The image enhancement parameter adjustment values of each unit are converted to obtain the spatial domain sigma value and gray-level domain sigma value of the line scan reflectance grayscale image, line scan depth grayscale image, and area array reflectance color image in the unit region. The sigma values of the three types of images are calculated according to the same rules to ensure processing consistency. Subsequently, these sigma values are used to dynamically configure the bilateral filter parameters of each image. The spatial domain sigma value is substituted into the spatial weight calculation module of the filter, and the gray-level domain sigma value is substituted into the gray-level weight calculation module. Each image is processed according to grid units. The filtering process is performed one by one. For each pixel in the image, the spatial weight and grayscale weight of all pixels in the neighborhood of that pixel are calculated according to the filter parameters of the cell. In a 3×3 neighborhood, the spatial weight is based on the distance between pixels and the spatial sigma value, and the grayscale weight is based on the pixel grayscale difference and the grayscale sigma value. The two weights are multiplied to obtain the comprehensive weight. Then, the grayscale values of the neighboring pixels are weighted and averaged using the comprehensive weight to obtain the filtered grayscale value of the pixel. After traversing all pixels, the noise-suppressed image data is obtained.
[0069] Step 312: Based on the preprocessing parameter adjustment values of the point cloud processing, calculate the dynamic distance threshold of the line-scan laser point cloud data; configure the parameters of the statistical filter using the dynamic distance threshold of the line-scan laser point cloud data; based on the parameters of the statistical filter, perform k-nearest neighbor-based statistical outlier filtering on the line-scan laser point cloud data to obtain denoised 3D point cloud data. Specifically, this includes: calculating the dynamic distance threshold of the line-scan laser point cloud data based on the point cloud processing preprocessing parameter adjustment values of each grid cell; firstly, determine the threshold calculation benchmark, based on the average distance deviation of the point cloud in the flat area of the steel plate surface (usually 0.05mm), and establish a conversion relationship between the parameter adjustment values and the distance threshold. The larger the parameter adjustment value, the smaller the dynamic distance threshold. For example, an adjustment value of 0.9 corresponds to a threshold of 0.08mm, and an adjustment value of 0.2 corresponds to a threshold of 0.3mm; according to this relationship, convert the point cloud processing parameter adjustment value of each cell into the dynamic distance threshold of that cell. Subsequently, the parameters of the statistical filter were configured using the dynamic distance threshold. The number of k-nearest neighbors for the statistical filter was set to 6, which is the optimal value verified in previous experiments. This means that the average distance between each point and its 6 nearest neighbors needs to be calculated. Based on the configured filter parameters, outlier filtering was performed on the line-scan laser point cloud data by grid cell. The point cloud data in each cell was traversed, and the Euclidean distance between each point and its 6 nearest neighbors was calculated and the average distance was calculated. If the average distance was greater than the dynamic distance threshold of the cell, the point was determined to be an outlier and removed. If it was less than or equal to the threshold, it was retained as a valid point. After traversing all cells, the denoised 3D point cloud data was obtained, ensuring that abnormal noise caused by water stains was removed while retaining real defect points.
[0070] Step 313: Perform contrast-enhancing grayscale stretching processing on the noise-suppressed image data to obtain enhanced image data; combine the enhanced image data with the denoised 3D point cloud data to form a preprocessed high-quality multi-source dataset, specifically including: firstly, performing contrast-enhancing grayscale stretching processing on the noise-suppressed line scan reflectance grayscale image, line scan depth grayscale image, and area array reflectance color image (after being converted to grayscale); for a single image, first count the grayscale values of all its pixels and find the minimum grayscale value G. min With the maximum grayscale value G max If G min With G max If the difference is too small, such as less than 50, then manually set G. min' =0、G max' =255 to expand the grayscale range; for the grayscale value G of each pixel in the image, press (GG min ) / (G max -G minThe stretched grayscale value G' is calculated using the method of 0 × 255. If G' is less than 0, it is set to 0; if it is greater than 255, it is set to 255. This ensures that the grayscale value of the stretched image is distributed within the range of 0 to 255, improving image contrast and texture detail clarity, resulting in enhanced image data. After image enhancement, the enhanced three types of image data are associated with the denoised 3D point cloud data according to a unified spatial coordinate system. That is, through the spatial correspondence established in step 204, it is ensured that each pixel of the image data accurately matches the corresponding 3D coordinate of the point cloud data, forming a preprocessed high-quality multi-source dataset, providing high-quality data input for subsequent cross-modal feature extraction of deep learning models.
[0071] In a preferred embodiment of the present invention, step 400 involves inputting the preprocessed high-quality multi-source dataset into a pre-trained multi-data fusion neural network to extract texture and geometric features, and then fusing them through a cross-modal attention mechanism to obtain the detection result, including:
[0072] Step 401: Input the enhanced image data from the preprocessed high-quality multi-source dataset into the 2D feature extraction branch of the pre-trained multi-data fusion neural network. Extract 2D feature vectors containing texture details using an improved ResNet-101 network. Simultaneously, input the denoised 3D point cloud data into the 3D feature extraction branch. Extract 3D feature vectors containing geometric shapes using a point cloud convolutional network. Specifically, before processing the input data, the construction and training of the pre-trained multi-data fusion neural network must be completed. The network construction aims to adapt to multi-source data feature extraction and cross-modal fusion. The specific construction process is as follows: A hierarchical design approach is adopted to build the overall network architecture, which includes a 2D feature extraction branch, a 3D feature extraction branch, a subsequent cross-modal attention fusion module, and an output branch. The 2D feature extraction branch is specifically used to process the texture features of the image data, and the 3D feature extraction branch is specifically used to process the geometric features of the point cloud data. The two are designed in parallel to ensure... To ensure simultaneous extraction of features from multiple data sources, for the 2D feature extraction branch, the ResNet-101 network was selected as the basic architecture and improved. At the output of each bottleneck block in the third convolutional stage (containing 23 bottleneck blocks) and the fourth convolutional stage (containing 36 bottleneck blocks), a Spatial Attention Mechanism (SAM) module was embedded in series. This SAM module consists of a global average pooling layer, a global max pooling layer, a 1×1 convolutional layer, and a Sigmoid activation function. The number of channels in the 1×1 convolutional layer was set to 1 / 4 of the number of output channels of the corresponding bottleneck block to achieve channel dimension compression and accurate generation of attention weights. For the 3D feature extraction branch, the PointNet network was improved by adding a neighborhood grouping module and a multi-scale convolutional layer. The neighborhood grouping module was implemented using the ball query algorithm. The multi-scale convolutional layer was set to 3 layers with channel dimensions of 64, 128, and 512 dimensions respectively. Each convolutional layer was followed by a batch normalization layer and a ReLU activation function to enhance the geometric feature extraction capability.
[0073] The training process of this network needs to be divided into two stages: pre-training and transfer learning fine-tuning. The specific training process is as follows: First, prepare the training dataset, using a hybrid dataset mode of public defect dataset and self-collected dataset. The public defect dataset uses the NEU-DET dataset, which contains 6 types of steel plate defects such as cracks and scratches, with a total of 3,000 samples. The self-collected dataset is obtained from the current production line and contains 10,000 steel plate defect samples of different specifications and surface conditions. Among them, the 10,000 samples of different specifications are 500 to 2,000 mm in width and 6 to 20 mm in thickness, and the different surface conditions are oxide scale and coating. Both datasets need to be labeled with the texture region of the defect (for images) and geometric coordinates (for point clouds). Then, data augmentation processing is performed on the dataset. For the image data, ±10° rotation, 0.7 to 1.3x scaling, and ±15% brightness adjustment are performed. For the point cloud data, ±5mm translation, ±10° local rotation, and random downsampling are performed, retaining 70% to 90% of the point cloud to improve the network performance. The generalization ability of the network was assessed. During the pre-training phase, the AdamW optimization algorithm was used, with a learning rate of 5e-5 and weight decay of 0.01. The 2D feature extraction branch used cross-entropy loss function to optimize texture feature extraction accuracy, and the 3D feature extraction branch used mean squared error loss function to optimize geometric feature extraction accuracy. The entire training process was iterated for 100 rounds. After each round, a validation set (20% of the total dataset) was used to evaluate the feature extraction performance. The pre-training phase ended when the 2D feature texture recognition accuracy on the validation set was greater than or equal to 94% and the 3D feature geometric representation error was less than or equal to 0.15mm. After pre-training, the network entered the transfer learning fine-tuning phase. 500 steel plate samples (including various defects) collected from the current production line were selected as the fine-tuning dataset. The learning rate was adjusted to 1e-5, and the training was iterated for 20 rounds. By fine-tuning the weight parameters of each layer of the network, the network was adapted to the steel plate features of the current production line. When the feature extraction error on the fine-tuning dataset was stably less than 10% of the error in the pre-training phase, the network training was completed, resulting in a pre-trained multi-data fusion neural network.
[0074] After completing network construction and training, input data processing and feature extraction can be carried out. First, the processing specifications for the input data are determined. The preprocessed high-quality multi-source dataset obtained in step 313, along with the enhanced image data, including line-scan reflectance grayscale images, line-scan depth grayscale images, and area array reflectance color images, undergoes a unified format conversion. The area array reflectance color images need to be converted to 3-channel grayscale images first. This conversion is achieved by weighted averaging of the RGB channel values, setting the R channel weight to 0.299, the G channel weight to 0.587, and the B channel weight to 0.114 to ensure... The converted grayscale image retains the texture information of the original color image; the line scan reflectance grayscale image and the line scan depth grayscale image are both single-channel formats, directly retaining single-channel features, ultimately enabling all three types of images to adapt to the input requirements of the 2D feature extraction branch in single-channel or 3-channel grayscale format; then the standardized image data is input into the 2D feature extraction branch in batches, with the batch size set to 8. This size is determined based on the parallel computing capabilities of the subsequent GPU (such as NVIDIA A100), ensuring efficient utilization of GPU resources while avoiding computational delays caused by excessively large batches.
[0075] The feature extraction process of the 2D feature extraction branch is as follows: Standardized image data is first input into the first convolutional stage of the improved ResNet-101 network. After processing with a 7×7 convolutional kernel (stride 2), batch normalization layers, and the ReLU activation function, an initial feature map with dimensions of 500×1600×64 is obtained. Subsequently, it passes through the second convolutional stage (3 bottleneck blocks, output feature map size 250×800×256) and the third convolutional stage (23 bottleneck blocks). After each bottleneck block output in the third convolutional stage, it is connected to the SAM module: The feature map output from the bottleneck block (size 250×800×256) is first subjected to channel-dimensional global average pooling and global max pooling to obtain... Two 1×1×256 feature vectors are concatenated into a 1×1×512 vector and then input into a 1×1 convolutional layer (64 output channels). The sigmoid activation function generates a 1×1×256 spatial attention weight map. This weight map is multiplied pixel-by-pixel with the original bottleneck block output feature map to enhance the feature response of key detail areas such as crack texture and oxide scale indentation texture. The feature map then enters the fourth convolutional stage (36 bottleneck blocks). In this stage, the SAM module is also connected after each bottleneck block output to further optimize the texture feature extraction effect. Finally, the feature map is processed by a global average pooling layer and then input into a fully connected layer, i.e., the output dimension is 1024, resulting in a two-dimensional feature vector containing texture details.
[0076] At the same time, the denoised 3D point cloud data (point density greater than or equal to 200 points / cm²) 2The input 3D feature extraction branch performs the following process: First, the input point cloud is randomly downsampled, fixing the number of points per frame to 1024. This number was determined experimentally to avoid both computational inefficiency due to an excessive number of points and geometric feature loss due to an insufficient number. Then, a ball query algorithm is used to group each downsampled point into neighborhoods, setting the group radius to 0.5mm and each group to contain 32 neighborhood points, ensuring sufficient capture of local geometric information for each point. The grouped point cloud features are then input into three stacked convolutional layers. Point cloud features consist of 3-dimensional coordinates and 3-dimensional normal vectors for each point, totaling 6 dimensions. The first convolutional layer maps the 6-dimensional features to 64-dimensional features, the second convolutional layer enhances the 64-dimensional features to 128-dimensional features, and the third convolutional layer further enhances the 128-dimensional features to 512-dimensional features. After each convolutional layer, a batch normalization layer is applied to eliminate feature distribution offset, and then a ReLU activation function is applied to enhance non-linear expressive power. Finally, the 512-dimensional local features output from the third convolutional layer are input into a global max pooling layer to aggregate the local features of all points, resulting in a 512-dimensional 3D feature vector containing geometric information such as pit depth distribution and convex contours.
[0077] Step 402: Based on the two-dimensional feature vector and the three-dimensional feature vector, channel dimension concatenation is performed to obtain the fused feature; the fused feature is input into the cross-modal attention mechanism to calculate the correlation weight between different modal features, and the fused feature is reconstructed based on the correlation weight to obtain the cross-modal fused feature vector. Specifically, the two-dimensional feature vector (1024-dimensional) and the three-dimensional feature vector (512-dimensional) output in step 401 are concatenated by channel dimension. Specifically, the channel dimension of the three-dimensional feature vector and the channel dimension of the two-dimensional feature vector are directly superimposed along the network feature channel axis (i.e., dimension axis) to form a fused feature with a total channel dimension of 1024 + 512 = 1536. This process ensures that the original information of the two types of modal features is included in the fusion process without omission, and avoids information loss caused by feature truncation.
[0078] The 1536-dimensional fused feature is then input into the cross-modal attention mechanism module. This module first splits the fused feature into channel dimensions, separating the 1024-dimensional channel part corresponding to the original 2D feature and the 512-dimensional channel part corresponding to the original 3D feature. Then, it calculates the correlation similarity between the two feature parts. Specifically, it calculates the inner product of the transpose matrices of the original 2D feature channel part and the original 3D feature channel part through matrix multiplication, resulting in a similarity matrix with dimensions of 1024×512. Each element in the matrix represents the correlation degree between a certain channel of the 2D feature and a certain channel of the 3D feature. Each row of the similarity matrix is then subjected to softmax normalization, converting the element values into correlation weights in the range of 0 to 1, with the sum of the weights in each row being 1. This weight quantitatively represents the supplementary importance of each channel of the 3D feature to the corresponding channel of the 2D feature.
[0079] Based on the aforementioned correlation weights, the fused features are reconstructed by weighting. Specifically, the feature value of each channel in the original two-dimensional feature channel is multiplied by the correlation weight of its corresponding row, and the feature value of each channel in the original three-dimensional feature channel is multiplied by the correlation weight of its corresponding column. The correlation weights are obtained by softmax normalization of the similarity matrix column direction. Then, the two weighted features are merged again along the channel axis to obtain a cross-modal fused feature vector with a dimension of 1536. This vector strengthens the key correlation features for defect identification, such as the correlation between crack texture and crack depth geometric features, weakens the interference of irrelevant features, and improves the feature's ability to represent composite defects through weight allocation.
[0080] Step 403: The cross-modal fusion feature vector is input to the defect classification branch and the regression branch. The softmax activation function of the defect classification branch outputs the defect category probability distribution, while the linear activation function of the regression branch outputs the defect location coordinates and three-dimensional size parameters. Finally, a detection result containing defect category, location information, and three-dimensional size is generated. Specifically, this includes: First, the 1536-dimensional cross-modal fusion feature vector obtained in step 402 is simultaneously input to the defect classification branch and the regression branch of the network to achieve parallel computation of the two types of outputs. The defect classification branch contains two fully connected layers. The first fully connected layer maps the 1536-dimensional features to 256-dimensional features through a weight matrix, and the second layer... The fully connected layer further maps the 256-dimensional features into 6-dimensional features, corresponding to 6 types of target defects: cracks, holes, pits, indented oxide scale, protrusions, and scratches. After the output of the second fully connected layer, a softmax activation function is applied. The core function of this softmax activation function is to convert the 6-dimensional feature values with no clear numerical range output by the second fully connected layer into non-negative probability values, and the sum of all probability values is 1. This transforms the abstract feature output into a probability expression of defect categories that conforms to probabilistic statistical logic. At the same time, by strengthening the weight of categories corresponding to high feature values and weakening the weight of categories corresponding to low feature values, the distinguishability between different defect categories is improved, adapting to the accurate classification requirements of the 6 types of target defects.
[0081] The specific transformation process involves converting each feature value to a non-negative value using a natural exponential function, and then dividing each non-negative value by the sum of the exponentially converted values of all six feature values to obtain the probability value for each type of defect. These generated probability values serve two core purposes: firstly, they act as a direct basis for preliminary defect category determination, selecting the category with the highest probability value as the preliminary defect category for the current region, ensuring clear quantitative support for the category determination; secondly, they serve as a core indicator for subsequent validity verification, providing a reference for judging whether there are valid defects in the current region and avoiding the one-sidedness caused by relying solely on feature output for category determination. Through the above processing, a defect category probability distribution is formed, and the category with the highest probability value is the preliminary defect category determination result for the current region.
[0082] Meanwhile, the regression branch also contains two fully connected layers. The first fully connected layer maps the 1536-dimensional fused features to 128-dimensional features, and the second fully connected layer maps the 128-dimensional features to 5-dimensional features. These 5-dimensional features correspond to the location coordinates and three-dimensional size parameters of the defect, namely the X-axis starting coordinate, the Y-axis starting coordinate, the defect length, the defect width, and the defect depth. The X-axis starting coordinate corresponds to the steel plate transmission direction, the Y-axis starting coordinate corresponds to the steel plate width direction, the defect length is the span in the X-axis direction, the defect width is the span in the Y-axis direction, and the defect depth is the deviation in the Z-axis direction. After the output of the second fully connected layer, a linear activation function is applied, and the output 5-dimensional feature values are directly retained as the final parameter results without nonlinear transformation to ensure the consistency of the parameter values with the actual physical dimensions. The parameter units are all in mm, consistent with the units of the unified spatial coordinate system.
[0083] Finally, the classification and regression results are validated. If the probability value of a certain type of defect is less than 0.5, it is determined that there are no valid defects in that area, and the corresponding regression parameter is removed. If the regression parameter exceeds the actual specification range of the steel plate, such as the defect length being greater than the width of the steel plate or the absolute value of the depth being greater than the thickness of the steel plate, the average parameter value of adjacent valid defects in the same batch of steel plates is retrieved for correction. After the validation is completed, the information is integrated in a structured format according to the defect category, location coordinates (X, Y), and three-dimensional dimensions (length, width, depth) to generate a detection result containing complete defect information. This result can be directly used for subsequent OPC UA protocol transmission and interaction with the production system.
[0084] In a preferred embodiment of the present invention, step 500, based on the detection results, generates a structured defect report, transmits it to the PLC system via the OPC UA protocol to trigger automatic marking and quality grading operations, and simultaneously stores the detection data in a blockchain database for quality traceability and model iteration, including:
[0085] Step 501: Based on the defect category, location information, and three-dimensional dimension parameters in the detection results, obtain a structured defect report containing defect feature descriptions, spatial coordinates, and geometric dimensions. Specifically, this includes: first, determining the core field system of the structured defect report. This system needs to cover the three core dimensions of defect features, spatial location, and geometric dimensions, and the field format needs to be compatible with subsequent OPC. The requirements for UA protocol transmission and PLC system parsing include seven fields: defect feature description, X-axis starting coordinate, Y-axis starting coordinate, defect length, defect width, defect depth, and detection timestamp. The defect feature description field must combine the defect category from the detection results with the corresponding key texture and geometric features to construct the field content. For example, a crack is defined as having a length of 2.5mm and a width of 0.3mm, extending along the X-axis, ensuring the description includes both category identification and detailed features. The spatial coordinate field extracts the X and Y axis coordinates of the defect location information from the detection results, and these coordinates must be consistent with the unified spatial coordinate system determined in step 205, retained to one decimal place in mm, to avoid subsequent positioning errors due to coordinate system deviation. The geometric dimension field extracts the three-dimensional dimension parameters from the detection results, also retained to one decimal place in mm, ensuring the length, width, and depth parameters correspond to the actual shape of the defect.
[0086] After extracting the above parameters, data format standardization is required. Defect feature descriptions are converted to string format (length less than or equal to 50 characters), spatial coordinates and geometric dimensions are converted to numerical format, and detection timestamps are converted to year-month-day hour:minute:second.millisecond format, accurate to 1 millisecond. Further data integrity is verified. If any field parameter is missing, such as if defect depth is not detected, it is marked as undetected and a note is added. For example, planar defects lack depth information to avoid null values causing report parsing failure. Finally, the fields are integrated in a fixed order: defect feature description, spatial coordinates (X, Y), geometric dimensions (length, width, depth), and detection timestamp, forming a structured defect report. This ensures that the report requires no additional format conversion during subsequent transmission and parsing, improving data flow efficiency.
[0087] Step 502: Transmit the structured defect report to the production line PLC control system via the OPC UA communication protocol, enabling the PLC system to obtain defect detection information. This includes: first, configuring the parameters of the OPC UA communication protocol to match the communication interface parameters of the production line PLC control system. This includes setting the communication port number, sampling period, and timeout. The communication port number is typically set to 4840, conforming to the OPC UA standard port specification. The sampling period is set according to the production line speed; for example, a sampling period of 100ms is set when the production line speed is 0.3m / s to ensure real-time performance. The timeout is set to 500ms to avoid transmission interruptions due to network latency. Then, establish the mapping relationship between the structured defect report fields and the PLC system variable nodes. For example, map the X-axis starting coordinate to the PLC's DefectX variable node, and the defect category to the DefectType variable node, ensuring that each report field has a unique corresponding PLC variable node to avoid data mapping confusion.
[0088] After the mapping relationship is established, the structured defect report is encapsulated using the OPC UA protocol. The data of each field in the report is converted into communication data packets according to the binary encoding format specified by the protocol. During the encapsulation process, a data check bit is added, i.e., the CRC32 check algorithm is used. The check value of the data packet is calculated to ensure that the data is not tampered with or lost during transmission. After encapsulation, the protocol communication link is started. A connection request is first sent to the PLC system. After receiving the connection confirmation signal from the PLC system, data packets are transmitted frame by frame according to the preset sampling period. During the transmission, the communication status is monitored in real time. If a transmission timeout or check failure occurs, a retransmission mechanism is immediately triggered. The number of retransmissions is less than or equal to 3. If the retransmission still fails, an alarm signal is sent to the system control unit to ensure that the PLC system can accurately and timely obtain defect detection information and provide data support for subsequent production operations.
[0089] Step 503: Based on the defect detection information, trigger the automatic marking device to perform inkjet marking at the defect location. Simultaneously, activate the quality grading module to determine the quality grade of the steel plate, obtaining production data containing marking and grading results. Specifically, this includes: first, processing the triggering logic of the automatic marking device. After receiving the defect detection information, the PLC system needs to determine the extension direction of the defect using the tangent equation algorithm of the parabola to optimize the adaptability of the marking geometric parameters. Specifically, first, extract the feature point coordinates of the defect along the surface extension direction from the detection results, and select the defect start point, midpoint, and end point. Three key feature points were identified, and their X and Y axis spatial coordinates were extracted to maintain consistency with a unified spatial coordinate system with an accuracy of 0.1 mm. For example, the starting point (1498.5, 800.2), midpoint (1500.0, 800.3), and ending point (1501.5, 800.4) were extracted. The coordinates of these three feature points were substituted into the parabolic fitting model, and the coefficients of the quadratic, linear, and constant terms of the parabola were adjusted using the least squares method to ensure that the deviation between the fitted curve and the coordinates of the three feature points was less than or equal to 0.1 mm. Finally, a parabolic equation that could characterize the extension trend of the defect was obtained.
[0090] Based on the parabolic equation, the tangent direction at the midpoint of the defect (i.e., the center point of subsequent marking) is further calculated. The X-axis coordinates of the midpoint are substituted into the calculation formula of the tangent equation of the parabola, and the tangent direction is determined by solving the slope value of the tangent. For example, if the slope is calculated to be 0.067, the angle between the tangent and the positive X-axis is about 3.8°, which is the direction of the defect's extension. This process ensures that the deviation between the tangent direction and the actual direction of the defect is less than or equal to 1°, providing a basis for the orientation design of the subsequent marking geometric parameters.
[0091] After determining the direction of defect extension, the X and Y axis spatial coordinates of the defect are converted into the mechanical coordinates of the automatic marking device. The conversion process requires consideration of the relative position parameters between the marking device and the detection area, including the X-axis offset and Y-axis offset of the marking device. These parameters are obtained through prior calibration with an accuracy of ±0.1mm. For example, if the X-axis spatial coordinate of the defect midpoint is 1500.0mm and the X-axis offset of the marking device is 50mm, then the mechanical X-coordinate is 1500.0 + 50 = 1550.0mm. The Y-axis coordinate is converted in the same way to ensure that the converted mechanical coordinates accurately correspond to the position of the defect midpoint.
[0092] Subsequently, the geometric parameters of the inkjet marking are calculated based on the defect length and tangent direction. An elliptical marking is used to adapt to the extended shape of the defect, rather than a fixed circular marking. If the defect length is less than or equal to 5mm, the major axis length of the ellipse is set to 3mm and the minor axis length to 2mm, with the major axis direction consistent with the tangent direction. If the defect length is greater than 5mm, the major axis length of the ellipse is set to 4mm and the minor axis length to 2mm, with the major axis direction also along the tangent direction. This parameter design ensures that the elliptical marking can completely cover the defect area, while avoiding excessive extension in the minor axis direction that could affect the surrounding normal steel plate surface. After the parameter calculation is completed, the PLC system sends a trigger signal to the automatic marking device. The control device's robotic arm adjusts the nozzle angle along the tangent direction, moves to the target mechanical coordinates, and performs the inkjet marking operation according to the set ellipse parameters. After marking is completed, the device sends a marking success signal back to the PLC system, simultaneously recording parameters such as the marking angle and major or minor axis dimensions.
[0093] Simultaneously, the quality grading module is activated to determine the quality level. This module's preset grading standards are determined based on the steel plate specifications (width, thickness) and the degree of defect impact. Specifically, a base score is set according to defect type, such as -30 points for cracks and -10 points for indented scale; additional scores are set according to defect size, such as -2 points for every 1mm increase in length and -5 points for every 0.1mm increase in depth; the base score for the steel plate is set at 100 points, and the total score = base score + base score + additional scores; the quality level is determined based on the total score: a total score of 90 or higher is considered excellent, 70 or lower is considered acceptable, and a total score less than 70 is considered substandard. Non-conforming products; during the grading process, the current steel plate specifications need to be retrieved from the production line MES system, such as width 1800mm and thickness 10mm, to ensure that the grading standard is compatible with different specifications of steel plates. For example, steel plates with a thickness of 8mm or less are more sensitive to the depth of the indentation, and the depth bonus is adjusted to add negative 8 points for every 0.1mm increase. After grading, the automatic marking results are integrated with the quality grade judgment results. The automatic marking results include: marking position, ellipse parameters, and angle. The quality grade judgment results include: grade and total score, forming production data containing timestamp, steel plate number, marking parameters, and grade parameters, which is fed back to the PLC system for temporary storage.
[0094] Step 504: Integrate the production data, inspection results, and multi-source data to form a complete inspection data package, store it in the blockchain database to establish a timestamped quality traceability record, and provide training samples for iterative updates of the deep learning model. Specifically, this includes: First, determining the structural framework of the complete inspection data package. This framework must include three core types of information: production data, inspection results, and multi-source data. The production data is the data generated in step 503, containing the labeling and grading results, retaining all fields. The inspection results are the original inspection data generated in step 403, containing the defect category, location, and size, retaining key parameters and removing redundant intermediate values. For the multi-source data, select preprocessed enhanced image keyframes (retaining image frames of the defect area and a surrounding 20mm range for each steel plate, with resolution compressed to 1000×800 to balance storage and clarity) and denoised point cloud keyframes (retaining point cloud data of the defect area, with point density reduced to 100 points / cm). 2 To reduce storage usage, ensure that data packets contain full-process information while avoiding excessive data redundancy.
[0095] During data integration, each steel plate is assigned a unique data association identifier, consisting of the steel plate's production batch number and inspection timestamp, such as Batch20240501-14:30:25.123. Production data, inspection results, and multi-source data are linked and bound through this identifier to form a complete inspection data package. Further, a timestamp and data hash value required for blockchain storage are generated. The timestamp is accurate to 1 millisecond and consistent with the inspection timestamp. The data hash value is calculated using the SHA-256 algorithm to ensure data uniqueness and immutability. Subsequently, the complete inspection data package, timestamp, and data hash value are encapsulated according to the blockchain database's write format requirements. A write request is submitted through the blockchain node's API interface. After receiving the request, the database verifies the validity of the data hash value. If verification is successful, the data is written to a block and synchronized to all nodes, establishing a timestamped quality traceability record to ensure data is searchable and verifiable during subsequent traceability processes.
[0096] Simultaneously, training samples suitable for deep learning model iteration are screened from the complete detection data package. The screening criteria are: clear defect category, complete 3D size parameters, and clear keyframes of multi-source data. Specifically, a clear defect category means a probability value greater than or equal to 0.8, complete 3D size parameters mean no undetected fields, and clear keyframes of multi-source data mean no blurring or noise. For data packages that meet the criteria, image blocks (200×200 pixels) and point cloud fragments (200 points) of the defect area are extracted, and pixel-level masks of the defects are added to supplement them. These are generated based on the position and size in the original detection results to form fully labeled training samples. More than 200 such new samples are collected each month and divided into training and validation sets in a 7:3 ratio. These samples are stored in the model iteration database to provide sample data that fits the actual production scenario for the online learning and updating of the deep learning model, helping the model to continuously adapt to the detection needs of the production line.
[0097] like Figure 2 As shown, embodiments of the present invention also provide a steel plate surface defect detection system based on multi-source data fusion, comprising:
[0098] The surface pretreatment module is used to pretreat the water stains on the steel plate surface, reduce the residual humidity to below a predetermined low humidity threshold, and obtain a dry surface.
[0099] The acquisition module is used to acquire multi-source data based on the dry surface through a synchronously triggered 3D laser line scan camera and a 2D industrial line scan camera, to obtain line scan laser point cloud data, line scan reflectivity grayscale image, line scan depth grayscale image and area array reflectivity color image, establish spatial correspondence between multi-source data, and form a multi-source dataset based on the spatial correspondence between multi-source data.
[0100] The processing module is used to establish a multi-source data calibration region on the surface of a steel plate based on a multi-source dataset and perform partitioning and meshing to obtain a mesh framework; analyze the multi-source data corresponding to the mesh framework to obtain preprocessing parameter adjustment values for image enhancement and point cloud filtering; use the preprocessing parameter adjustment values to perform adaptive bilateral filtering and grayscale stretching on the line scan reflectivity grayscale image, the line scan depth grayscale image, and the area array reflectivity color image, and simultaneously perform outlier removal on the line scan laser point cloud data to obtain a high-quality preprocessed multi-source dataset;
[0101] The defect analysis module is used to input the pre-processed high-quality multi-source dataset into a pre-trained multi-data fusion neural network, extract texture and geometric features, and obtain the detection results after fusion through a cross-modal attention mechanism;
[0102] The results feedback module is used to generate a structured defect report based on the detection results, which is transmitted to the PLC system via the OPC UA protocol to trigger automatic marking and quality grading operations. At the same time, the detection data is stored in the blockchain database for quality traceability and model iteration.
[0103] It should be noted that this system corresponds to the method described above, and all implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect. The above description represents the preferred embodiment of the present invention. It should be pointed out that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for detecting surface defects in steel plates based on multi-source data fusion, characterized in that, The method includes: Step 100: Perform water stain pretreatment on the surface of the steel plate to reduce the residual humidity to below the predetermined low humidity threshold and obtain a dry surface; Step 200: Based on the dry surface, multi-source data is collected by synchronously triggered 3D laser line scan camera and 2D industrial line scan camera to obtain line scan laser point cloud data, line scan reflectivity grayscale image, line scan depth grayscale image and area array reflectivity color image, establish spatial correspondence between multi-source data, and form multi-source dataset based on spatial correspondence between multi-source data. Step 300: Based on the multi-source dataset, a multi-source data calibration region is established on the steel plate surface and partitioned into a grid to obtain a grid framework; the multi-source data corresponding to the grid framework is analyzed to obtain preprocessing parameter adjustment values for image enhancement and point cloud filtering; using the preprocessing parameter adjustment values, adaptive bilateral filtering and grayscale stretching are performed on the line-scanned reflectivity grayscale image, the line-scanned depth grayscale image, and the area array reflectivity color image, while outlier removal is performed on the line-scanned laser point cloud data to obtain a preprocessed high-quality multi-source dataset, including: Based on the aforementioned grid framework, the density distribution characteristics of point cloud data and the gray-level gradient characteristics of image data within each grid cell are calculated respectively. The density distribution characteristics of the point cloud data are normalized to obtain normalized density coefficients, and the gray-level gradient characteristics of the image data are statistically analyzed to obtain the average gradient intensity. The normalized density coefficient and the average gradient intensity are weighted and fused according to a predetermined weight ratio to obtain a weighted fusion result. Based on the weighted fusion results, preprocessing parameter adjustment values for image enhancement and preprocessing parameter adjustment values for point cloud processing are generated simultaneously. Step 400: Input the preprocessed high-quality multi-source dataset into the pre-trained multi-data fusion neural network to extract texture and geometric features, and then fuse them through a cross-modal attention mechanism to obtain the detection results; Step 500: Based on the detection results, a structured defect report is obtained and transmitted to the PLC system via the OPC UA protocol to trigger automatic marking and quality grading operations. At the same time, the detection data is stored in the blockchain database for quality traceability and model iteration.
2. The method for detecting surface defects in steel plates based on multi-source data fusion according to claim 1, characterized in that, Step 100 includes: The surface humidity of the steel plate entering the detection area is monitored by an infrared humidity sensor to obtain the surface humidity status signal. The surface humidity status signal is processed, and when liquid water interference is detected, a water removal device start signal is obtained. According to the start signal of the dewatering device, the high-pressure air blowing unit and the mechanical scraping unit are controlled to perform a coordinated dewatering operation to obtain the steel plate surface after preliminary dewatering. The surface of the steel plate after preliminary dehydration is monitored for humidity in real time. When the monitoring data indicates that the residual humidity has reached the predetermined low humidity threshold, an operation termination signal is obtained and it is confirmed that a dry surface has been obtained.
3. The method for detecting surface defects of steel plates based on multi-source data fusion according to claim 2, characterized in that, Step 200 includes: Based on the dry surface, a 3D laser line scanning camera and a 2D industrial line scanning camera are controlled by a synchronous trigger signal to acquire raw 3D sensing data and 2D image data. The original 3D sensing data is reconstructed into point cloud and analyzed for reflectivity to obtain line scan laser point cloud data, line scan reflectivity grayscale image and line scan depth grayscale image; The 2D image data is subjected to color space conversion and pixel calibration to obtain an area array reflectance color image; Based on the line-scanned laser point cloud data and the area array reflectivity color image, a spatial correspondence between the line-scanned laser point cloud data and the area array reflectivity color image is established through a pre-calibrated sensor pose transformation matrix. Based on the spatial correspondence, and utilizing the inherent index association between the line scan reflectivity grayscale image, the line scan depth grayscale image, and the line scan laser point cloud data, all multi-source data are unified under the same spatial coordinate system to form a multi-source dataset with a unified spatial correspondence.
4. The method for detecting surface defects of steel plates based on multi-source data fusion according to claim 3, characterized in that, Step 300 includes: Based on the unified spatial coordinate system of the multi-source dataset, the boundary coordinates of the effective detection area on the steel plate surface are extracted; the geometric center position and normal vector direction of the detection area are calculated based on the boundary coordinates of the effective detection area on the steel plate surface. A local coordinate system for the steel plate surface is established based on the geometric center position and the direction of the normal vector; the boundary coordinates of the scanning coverage area are calculated in the local coordinate system according to the horizontal and vertical scanning field of view of the 3D laser line scanning camera. The spatial range of the multi-source data calibration area on the steel plate surface is determined based on the boundary coordinates of the scan coverage area, and the multi-source data calibration area is established. Based on the spatial range of the multi-source data calibration area, the reference grid cell size that meets the requirements of point cloud data feature analysis is calculated; Based on the reference grid unit size, the number of grid divisions is calculated along the longitudinal and transverse directions of the steel plate surface, respectively. Based on the number of grid divisions, uniformly distributed regular geometric units are established within the multi-source data calibration region to obtain a grid framework that perfectly matches the spatial range of the multi-source data calibration region.
5. The method for detecting surface defects in steel plates based on multi-source data fusion according to claim 4, characterized in that, Step 300 further includes: Based on the preprocessing parameter adjustment values of the image enhancement processing, the spatial domain sigma value and gray domain sigma value corresponding to the line scan reflectance grayscale image, the line scan depth grayscale image, and the area array reflectance color image are calculated respectively; the bilateral filter parameters of each image are dynamically configured using the spatial domain sigma value and the gray domain sigma value; filtering processing is performed on each image using the configured bilateral filter parameters to obtain image data after noise suppression; Based on the preprocessing parameter adjustment values of the point cloud processing, the dynamic distance threshold of the line-scan laser point cloud data is calculated; the parameters of the statistical filter are configured using the dynamic distance threshold of the line-scan laser point cloud data; based on the parameters of the statistical filter, k-nearest neighbor-based statistical outlier filtering is performed on the line-scan laser point cloud data to obtain denoised 3D point cloud data. The noise-suppressed image data is subjected to contrast-enhancing grayscale stretching to obtain enhanced image data; the enhanced image data is combined with the denoised 3D point cloud data to form a preprocessed high-quality multi-source dataset.
6. The method for detecting surface defects of steel plates based on multi-source data fusion according to claim 5, characterized in that, Step 400 includes: The enhanced image data from the preprocessed high-quality multi-source dataset is input into the 2D feature extraction branch of the pre-trained multi-data fusion neural network. The improved ResNet-101 network extracts two-dimensional feature vectors containing texture details. At the same time, the denoised three-dimensional point cloud data is input into the 3D feature extraction branch. The point cloud convolutional network extracts three-dimensional feature vectors containing geometric shapes. Based on the two-dimensional and three-dimensional feature vectors, channel dimensions are concatenated to obtain fused features; the fused features are input into a cross-modal attention mechanism to calculate the correlation weights between features of different modalities; the fused features are then reconstructed based on the correlation weights to obtain a cross-modal fused feature vector. The cross-modal fusion feature vector is input into the defect classification branch and the regression branch. The softmax activation function of the defect classification branch outputs the probability distribution of the defect category, while the linear activation function of the regression branch outputs the defect location coordinates and three-dimensional size parameters. Finally, a detection result containing defect category, location information and three-dimensional size is generated.
7. The method for detecting surface defects in steel plates based on multi-source data fusion according to claim 6, characterized in that, Step 500 includes: Based on the defect category, location information, and three-dimensional size parameters in the detection results, a structured defect report containing defect feature description, spatial coordinates, and geometric dimensions is obtained. The structured defect report is transmitted to the production line PLC control system via the OPC UA communication protocol, so that the PLC system can obtain defect detection information. Based on the defect detection information, the automatic marking device is triggered to perform inkjet marking at the defect location, and the quality grading module is activated to determine the quality grade of the steel plate, thereby obtaining production data containing the marking and grading results. The production data, test results, and multi-source data are integrated to form a complete test data package, which is stored in a blockchain database to establish a timestamped quality traceability record and provide training samples for iterative updates to the deep learning model.
8. A steel plate surface defect detection system based on multi-source data fusion, wherein the system implements the method as described in any one of claims 1 to 7, characterized in that, include: The surface pretreatment module is used to pretreat the water stains on the steel plate surface, reduce the residual humidity to below a predetermined low humidity threshold, and obtain a dry surface. The acquisition module is used to acquire multi-source data based on the dry surface through a synchronously triggered 3D laser line scan camera and a 2D industrial line scan camera, to obtain line scan laser point cloud data, line scan reflectivity grayscale image, line scan depth grayscale image and area array reflectivity color image, establish spatial correspondence between multi-source data, and form a multi-source dataset based on the spatial correspondence between multi-source data. The processing module is used to establish a multi-source data calibration region on the surface of a steel plate based on a multi-source dataset and perform partitioning and meshing to obtain a mesh framework; analyze the multi-source data corresponding to the mesh framework to obtain preprocessing parameter adjustment values for image enhancement and point cloud filtering; Using the preprocessing parameter adjustment values, adaptive bilateral filtering and grayscale stretching are performed on the line scan reflectivity grayscale image, the line scan depth grayscale image, and the area array reflectivity color image. At the same time, outlier removal is performed on the line scan laser point cloud data to obtain a high-quality multi-source dataset after preprocessing. The defect analysis module is used to input the pre-processed high-quality multi-source dataset into a pre-trained multi-data fusion neural network, extract texture and geometric features, and obtain the detection results after fusion through a cross-modal attention mechanism; The results feedback module is used to generate a structured defect report based on the detection results, which is transmitted to the PLC system via the OPC UA protocol to trigger automatic marking and quality grading operations. At the same time, the detection data is stored in the blockchain database for quality traceability and model iteration.
9. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 7.