A steel plate surface defect recognition method fusing visual attention mechanism
By aligning image sequences from linear and area array cameras and modeling two-dimensional Gaussian irradiance, combined with visual attention feature encoding and graph attention networks, the problem of inaccurate identification in steel plate surface defect recognition was solved, achieving real-time accurate identification and intelligent decision-making, thus improving identification accuracy and system reliability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- RIZHAO YULAN NEW MATERIAL CO LTD
- Filing Date
- 2026-03-24
- Publication Date
- 2026-06-23
AI Technical Summary
Existing technologies for identifying defects on steel plate surfaces suffer from problems such as high labor intensity in manual judgment, poor robustness of identification, and lack of modeling of the topological correlation and evolution logic of defect distribution, resulting in inaccurate identification and difficulty in adapting to changing production environments.
Image sequences are captured using linear and area array cameras. A standardized tensor is constructed through subpixel-level spatial alignment and two-dimensional Gaussian irradiance modeling. Combined with visual attention feature encoding network and graph attention network, the correlation modeling and topological reasoning between defects and process parameters are realized, the interference of ambient light is reduced, the process contribution weight is quantified, and a process target evaluation baseline is generated.
It enables real-time and accurate identification and intelligent decision-making for steel plate surface defects, improving identification accuracy and system decision reliability, adapting to complex production environments, and reducing false alarms and missed detections.
Smart Images

Figure CN122265208A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of defect detection technology based on computer vision, and particularly relates to a method for identifying defects on the surface of steel plates by incorporating a visual attention mechanism. Background Technology
[0002] As a core raw material in modern manufacturing, the surface quality of steel plates directly determines the structural strength and service life of subsequently processed parts. During the rolling, continuous casting, and conveying processes of steel plates, the complex physical environment makes the surface highly susceptible to defects such as cracks, folds, inclusions, and scratches. If these surface defects cannot be accurately identified and located in real time, substandard products will flow into downstream processes, leading to significant quality risks and economic losses. Therefore, achieving continuous automatic monitoring of the surface condition of steel plates on high-speed production lines is a core means of guiding quality improvement and enhancing the automation level of production lines.
[0003] Currently, the industry mainly uses the following methods to identify surface defects in steel plates: (1) Manual visual observation method. This method mainly relies on quality inspectors to obtain defect characteristics through observation. Since the production environment is often accompanied by high temperature, dust and vibration interference, and the production line moves at a fast speed, manual judgment is not only extremely labor-intensive, but the judgment results are also significantly affected by subjective experience and physiological fatigue. This makes it difficult to maintain the same inspection standards across different shifts, and there are physical limitations in the ability to capture subtle defects; (2) Conventional digital image processing methods. These methods mainly use fixed convolution kernels or morphological operators to extract edge features. Due to the complex mechanical processing texture on the surface of steel plates and the fact that changes in the incident angle of the light source can easily produce non-uniform reflections, conventional algorithms have poor robustness in the face of defects such as low contrast or strong background interference, and are difficult to adapt to the changing production environment; (3) Recognition methods based on general deep learning. Existing algorithms mostly focus on extracting local pixel features, lacking modeling of the topological relationships and evolution logic of defect distribution. In addition, most models fail to effectively integrate the physical laws of reflectivity consistency in metal surface imaging, making them prone to false alarms when faced with complex noise or non-stationary lighting. At the same time, most existing technologies only output static classification results, lacking the ability to perform strategic evaluation and automated processing of recognition results. Summary of the Invention
[0004] To address the above problems, this invention proposes a method for identifying surface defects in steel plates that integrates a visual attention mechanism, comprising the following steps: S1: Use linear and area array cameras to capture linear and area array image sequences of the steel plate surface, obtain the absolute physical coordinates of the production line encoder, perform spatiotemporal sequence alignment to obtain an aligned and normalized feature map sequence; use a two-dimensional Gaussian distribution function to establish an ambient light intensity mapping map and perform channel stitching to output a standardized surface state feature tensor with a defined dimension. Simultaneously acquire process feature sets and historical high-quality sample sets; S2, construct a visual attention feature encoding network that integrates physical information constraints, remove the influence of ambient light brightness gain on image quality, and obtain a high-dimensional physical-visual joint feature vector; S3, based on the absolute physical coordinates of the production line encoder and the joint feature vector of high-dimensional physical vision, performs defect and process heterogeneous data association modeling, quantifies the contribution weight of process parameters to defect formation through dual-track machine learning algorithm, obtains the current process feature vector and process association mapping matrix, and introduces process feature set and historical high-quality sample set to construct process target evaluation baseline. S4. Construct a steel plate surface defect detection network, and input the high-dimensional physical vision joint feature vector, the current process feature vector, the process correlation mapping matrix and the process target evaluation baseline into the network to obtain the defect detection results.
[0005] Preferably, S1 specifically includes: First, multi-source sensor deployment and raw data acquisition are performed simultaneously. Linear array and area array cameras are used to capture linear array image sequences of the steel plate surface. With array image sequence And bind the absolute physical coordinates of the production line encoder. With process feature set Secondly, spatiotemporal sequence alignment and standardization are performed, utilizing the absolute displacement values of the production line. and Calculate the relative offset in physical space s and pixel offset value The spatial deviation between the two images is eliminated by subpixel-level interpolation resampling, and an aligned and normalized feature map sequence is output. Finally, a standardized surface state feature tensor is constructed, and an ambient light intensity mapping map is established using a two-dimensional Gaussian distribution function. It then performs channel splicing and outputs a standardized surface state feature tensor with defined dimensions. ; The process feature set This includes the temperature of the heating furnace soaking zone, the tension fluctuation value between stands, the rate of change of the rolling exit speed, and the water pressure distribution parameters of the laminar cooling zone. The number of process parameters; the historical high-quality sample set is formed by screening the quality inspection records corresponding to the steel plate samples in the quality inspection system during synchronous reading.
[0006] Preferably, the spatiotemporal sequence alignment and standardization process is as follows: Linear image sequence The array image sequence and the absolute displacement values of the production line recorded at the corresponding sampling time. and Input data and set the camera reference mounting distance to a fixed value of 500mm; Secondly, calculate the physical spatial relative offset between linear and area array images. ; Then, the physical space relative offset Divide by a fixed value of the physical size of a unit pixel to calculate the pixel offset value. Linear image sequence Using the coordinates as a reference, a bicubic interpolation algorithm is employed, utilizing the pixel offset values. Opponent image sequence Resampling is performed to obtain a spatially aligned area image sequence. ;Will Spatial aligned area array image sequence Perform uniform scaling and normalization; Finally, the normalized linear image sequence and the area image sequence are concatenated along the channel dimension to output an aligned and normalized feature map sequence. ;in, The linear array image sequence of the first channel is denoted as The second channel's array image sequence is denoted as .
[0007] Preferably, the step of establishing an ambient light intensity mapping map using a two-dimensional Gaussian distribution function and performing channel stitching to output a standardized surface state feature tensor with determined dimensions specifically involves: Based on the absolute displacement value of the production line The spatial position of the current image frame on the production line is determined, and the irradiance value of each pixel within the sampling area is calculated using a two-dimensional Gaussian distribution function. The resulting image is a mapping map with the dimension of ambient light intensity. ; Align and normalize the feature map sequence Mapping with ambient light intensity Perform feature concatenation operations along the channel dimension to construct a standardized surface state feature tensor. And output; wherein, the first channel stores linear array enhanced texture features. Geometric features of the second channel storage array Third channel storage environment light intensity mapping map .
[0008] Preferably, the visual attention feature encoding network in S2 is specifically: First, visual attention block segmentation and linear embedding are performed to normalize the surface state feature tensor. Projection as block embedding matrix And superimpose the position encoding matrix Generate the initial input sequence Secondly, multi-head self-attention feature encoding is performed for the initial input sequence. Establish long-range feature dependencies and generate preliminary spatial feature maps using Transformer encoding blocks. ; Reconstruct the physical reflection component prediction branch, and use the initial spatial feature map Input into this branch and obtain the predicted intrinsic surface reflectance value. Subsequently, the physical information neural network PINN is introduced to predict the intrinsic reflectivity of the surface. Constraint correction is performed to reduce the irradiance of ambient light sources using physical imaging theory. The resulting impact on brightness gain, and the calculation of theoretically predicted intensity values. With image observation intensity Physical constraint loss function between Finally, the physical eigenvalue map is obtained based on physical consistency correction. and the preliminary spatial feature map Deep concatenation generates high-dimensional physical-visual joint feature vectors .
[0009] Preferably, the specific process of S3 is as follows: First, perform spatiotemporal alignment of defect and process heterogeneous data, using the absolute physical coordinates of the production line encoder obtained from S1. Perform an inverse mapping from the spatial domain to the temporal domain to obtain the defect feature vector. , defect feature vector Joint feature vectors with high-dimensional physical vision Perform splicing defect process associated state tensor Secondly, perform correlation reasoning based on a dual-track machine learning algorithm to analyze the defect process correlation state tensor. Process feature vectors are obtained by performing process feature channel extraction and average pooling operations. Using the regression coefficient matrix as Elastic network regression model and ensemble learning model Predict the continuous geometric index vector respectively With the defect category probability vector ; Perform process contribution quantification extraction again, combined with the regression coefficient matrix Compared with defect category probability vectors Attribution analysis to obtain the process correlation mapping matrix Finally, based on the historical high-quality sample set output from S1... Process feature set process feature vector and process correlation mapping matrix Calculate the historical optimal process configuration Incremental prediction results of defect traits Generate a baseline for evaluating process objectives. .
[0010] Preferably, the execution defects and heterogeneous process data are spatiotemporally aligned, specifically as follows: First, using the absolute physical coordinates of the production line encoder Perform the inverse mapping from the spatial domain to the temporal domain: Let the physical coordinates corresponding to the current defect feature be... In the absolute physical coordinates of the production line encoder Chinese query and Encoder coordinates with minimum Euclidean distance According to the production line operating speed and the inherent sampling delay of the sensor Determine the time window for process parameters that are strongly correlated with defect generation; the time window is defined as an interval of... to( ;in for The corresponding timestamp, The time delay tolerance is determined based on the thermal conductivity and plastic deformation characteristics of the steel plate; Secondly, within the time window to( Internally, a mean statistical operation is performed on each process parameter to obtain the process feature vector corresponding to the current defect feature. Finally, the defect feature vector Joint feature vectors with high-dimensional physical vision Perform one-to-one pairing and splice along the channel dimension to generate a defect process-related state tensor. .
[0011] Preferably, the process of performing associative reasoning based on a dual-track machine learning algorithm is as follows: From the defect process associated state tensor Extracting the process features corresponding to the post Each channel is used to perform average pooling along the spatial dimension to obtain the current process feature vector. Secondly, define the output attributes of the defect data: a vector of continuous geometric indices of the defects. Includes vertical length, horizontal width, and prediction depth; defect category probability vector. The dimension is 4, and its four components correspond to the predicted probabilities of four types of defects: cracks, folds, inclusions and scratches. Elastic network regression inference for continuous defect geometric indices: incorporating process feature vectors Input an elastic network regression model and output the prediction results of a continuous geometric index. The regression coefficient matrix corresponding to the elastic network regression model is denoted as... ;in, Columns 1, 2, and 3 represent the mapping relationship between each process parameter and the longitudinal length, transverse width, and predicted depth, respectively. Integrated classification reasoning based on defect category identification results: integrating process feature vectors Input ensemble learning model Output defect probability vector The ensemble learning model It consists of 5 regression tree models.
[0012] Preferably, the quantitative extraction of the execution process contribution is carried out as follows: First, the regression coefficient matrix The absolute values of each column element are taken, and an average operation is performed along the continuity geometric index dimension to obtain the process contribution weight vector corresponding to the continuity defect. ;in, The The component represents the first... The combined linear contribution weight of each process parameter to the geometric indices of continuous defects; Secondly, SHAP is used to analyze the defect category probability vector. Attribution analysis is performed to obtain the process contribution weights corresponding to discrete defects; for the current process feature vector The first in Each process parameter, integrated learning model Using the complete process feature vector as input, the predicted probability vectors of each type of defect are obtained through inference. ; then remove the first Each process parameter is again based on an ensemble learning model. Perform model inference to obtain the predicted probability vectors for each type of defect. ; Calculate the difference between the two predicted probability vectors to obtain the first... The contribution vector of each process parameter to the prediction of each type of defect in this sample. ; For contribution vector The four components are averaged to obtain the first... The scalar contribution of each process parameter to the prediction of the discrete defect category of this sample; perform this inference operation on all samples and take the average to obtain the first... Discrete contribution weights corresponding to each process parameter ; for all Repeat the above process for each process parameter to obtain the process contribution weight corresponding to discrete defects. ; Finally, the process contribution weight vector corresponding to the continuity defect. Process contribution weight vector corresponding to discrete defects Perform element-wise averaging to obtain the overall process-related weight vector. Then, the overall process associated weight vector is used. Construct a diagonal process association mapping matrix .
[0013] Preferably, the baseline for evaluating the generation process target is defined as follows: First, calculate the historical high-quality sample set. Each sample corresponds to a process feature set The historical optimal process configuration is obtained by statistically averaging the parameters of each dimension. .
[0014] Subsequently, the process feature vector is calculated. Compared to the historical optimal process configuration Deviation vector Next, matrix multiplication is performed. Obtain the prediction results of the defect trait increment, which reflects the change in defect severity. .
[0015] Finally, based on the aforementioned historical optimal process configuration Incremental prediction results of defect traits and predetermined fixed weighting coefficients Generate a baseline for evaluating process objectives. Among them, the weighting coefficient Optimization based on historical data: Collect production data for 30 consecutive days, with the optimization objective being to minimize the actual defect rate. A grid search method is used to traverse the interval [0,1] with a step size of 0.1. The selection process ultimately yields the result that minimizes the sum of the defect false positive rate and the defect missed detection rate on the validation set. The value is used as a fixed parameter of the system.
[0016] Preferably, S4 specifically includes: First, a defect topology graph is constructed, and the high-dimensional physical-visual joint feature vector output by S2 is processed. The defect probability density map is obtained through upsampling and global average pooling operations. and visual node feature matrix Then, using the defect probability density map Candidate region nodes are extracted and an initial topology graph is constructed; then, process constraint feature mapping and joint node state construction are performed, and the process association mapping matrix output by S3 is applied. With process feature vector Perform a linear mapping to obtain the process risk characterization vector. The baseline for evaluating process objectives output by S3. Adding them together yields the process constraint description vector. Describing the process constraints as vectors With visual node feature matrix The visual-process joint node feature matrix is obtained by splicing. Finally, the initial topology graph and the vision-process joint node feature matrix are combined. The input graph attention network performs joint topological inference, and outputs a topology-enhanced feature matrix through two layers of graph attention propagation. Topology-enhanced feature matrix Input multi-task output head to obtain defect location mask Defect category identification results The defect categories include four types: cracks, folds, inclusions, and scratches.
[0017] Compared with the prior art, the present invention has the following innovative features and beneficial effects: (1) Design of a standardized tensor construction system integrating multi-source image physical alignment and ambient lighting modeling. A data benchmark construction scheme based on sub-pixel spatial alignment and two-dimensional Gaussian irradiance modeling is proposed. The system first deploys a joint imaging unit composed of linear and area industrial cameras to simultaneously acquire linear image sequences reflecting microscopic details and area image sequences reflecting macroscopic geometric features. Then, the physical spatial relative offset is calculated using the absolute displacement value of the production line encoder, and sub-pixel spatial alignment of the two images is achieved through bicubic interpolation algorithm. Most importantly, an ambient lighting intensity mapping map based on two-dimensional Gaussian distribution is introduced. The distribution characteristics of light source with bright center and dark edge in the imaging field of view are simulated using physical function relationship. This map is then deeply stitched with the image channel to construct a standardized surface state feature tensor with a defined dimension, thus providing a standardized input with highly aligned physical properties for subsequent models. (2) Design of visual attention feature encoding strategy with PINN constraint embedded in physical information neural network. An end-to-end feature extraction strategy based on physical imaging law constraint is proposed. In the feature encoding process, firstly, the multi-head attention layer of the Transformer encoding block is used to establish long-distance feature dependencies across image blocks and cross-viewpoint channels to generate a preliminary spatial feature map; then, a physical reflectance component prediction branch is constructed, and the feature space is mapped to the surface intrinsic reflectance prediction value through a three-layer continuous convolution sequence; by introducing the intensity value predicted by physical imaging theory and the real-time ambient light irradiance data, a self-supervised feedback closed loop containing a physical constraint loss function is constructed; this mechanism uses the physical consistency residual gradient backpropagation to drive the model to correct the reflectance prediction deviation in a directional manner, effectively weakening the brightness gain interference caused by non-stationary illumination, and ensuring that the generated physical intrinsic feature map can accurately anchor the real physical response of the steel plate surface; (3) Defect topology reasoning system based on heterogeneous data spatiotemporal correlation mapping and graph attention network. To address the difficulties in corresponding defect features with production line processes and the lack of topological correlation, a full-link process correlation inversion and topology reasoning system is established. First, the absolute physical coordinates of the production line encoder are used to perform the inverse mapping from the spatial domain to the temporal domain. Combined with the thermal conduction and plastic deformation characteristics of the steel plate, the time window of the process parameters is established to achieve accurate alignment between defect features and process feature vectors. Second, the process contribution weight is extracted using elastic network regression and ensemble learning models. The incremental prediction results of defect characteristics are calculated by matrix multiplication and a process target evaluation baseline is generated. Most importantly, a joint reasoning module based on graph attention network (GAT) is constructed. With the vision-process joint node feature matrix as input, the weighted aggregation of local texture and process constraint information is performed in the graph domain space, realizing a complete logical closed loop from macroscopic performance indicators to microscopic defect topology and then to classification decision. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the following description is only one embodiment of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating the overall technical route of the present invention.
[0020] Figure 2 This is a diagram of the visual attention feature encoding network framework of the present invention.
[0021] Figure 3 This is the topological association reasoning diagram of the present invention.
[0022] Figure 4This is a comparison test result diagram of the performance of multiple algorithms in steel plate surface defect identification in the embodiments of the present invention.
[0023] Figure 5 The figure shows the results of the anti-reflective interference ablation experiment before and after the introduction of physical constraints in the embodiment.
[0024] Figure 6 This is a correlation curve of the recognition distribution based on the candidate region extraction threshold in the embodiment. Detailed Implementation
[0025] To achieve real-time and accurate identification and intelligent decision-making for steel plate surface defects, this invention proposes a steel plate surface defect identification method that integrates a visual attention mechanism. The overall process is as follows: Figure 1 As shown, this method establishes a standardized benchmark with aligned physical properties and unified data modalities by simultaneously acquiring multi-source images of steel plates, production line process data, and quality inspection data, and constructing standardized tensors. Based on the Transformer encoding architecture, a visual attention feature encoding network integrating physical information constraints is built. The PINN constraint correction mechanism of the physical information neural network is used to remove the influence of ambient light gain on image quality, achieving accurate extraction of intrinsic reflective property features of defects. Furthermore, a spatiotemporal collaborative alignment and correlation modeling strategy for heterogeneous defect and process data is implemented. A dual-track machine learning algorithm quantifies the contribution weight of process parameters to defect formation, constructing a process target evaluation baseline. Finally, defect topology reasoning and multi-task prediction based on the Graph Attention Network (GAT) are executed. A joint node feature matrix is used to achieve deep fusion of local visual features and global process constraints, thereby outputting defect location masks and category recognition results, significantly improving the accuracy of steel plate quality inspection and the reliability of system decisions.
[0026] The implementation process of the present invention will be further described below with reference to specific embodiments.
[0027] S1. Data Synchronization Acquisition and Standardized Tensor Construction This step aims to construct a standardized baseline feature tensor with aligned physical properties and unified data modalities. This tensor establishes a strict physical mapping relationship between steel plates from different batches across four dimensions: surface microtexture, macroscopic three-dimensional geometric features, real-time production line process parameters, and historical quality inspection records. This provides definite data input for subsequent deep feature encoding and process correlation analysis. The overall process is as follows: First, multi-source sensor deployment and synchronous acquisition of raw data are performed. Linear array and area array cameras are used to capture linear array image sequences of the steel plate surface. With array image sequence And bind the absolute physical coordinates of the production line encoder. With process feature set At the same time, a high-quality historical sample set was formed based on the GB / T3274-2017 standard. Secondly, spatiotemporal sequence alignment and standardization are performed, utilizing the absolute displacement values of the production line. and Calculate the relative offset in physical space s and pixel offset value The spatial deviation between the two images is eliminated by subpixel-level interpolation resampling, and an aligned and normalized feature map sequence is output. Finally, a standardized surface state feature tensor is constructed, and an ambient light intensity mapping map is established using a two-dimensional Gaussian distribution function. It then performs channel splicing and outputs a standardized surface state feature tensor with defined dimensions. .
[0028] S1-1 Multi-Source Sensor Deployment and Synchronous Acquisition of Raw Data. First, a joint imaging unit is deployed along the steel plate's travel direction in the terminal area of the continuous steel plate conveying production line. This unit consists of one linear scan industrial camera and one area scan industrial camera. The linear scan camera is used to acquire high-resolution surface texture, and the area scan camera is used to capture three-dimensional geometric features. Synchronous sampling is performed at a fixed frame rate of 30fps under the same trigger clock. The linear scan industrial camera outputs a real-time sequence of single-channel grayscale linear scan images reflecting the surface's microscopic details. The industrial camera outputs a single-channel grayscale area array image sequence in real time, reflecting macroscopic three-dimensional features. Simultaneously, at each synchronous sampling moment, the absolute displacement value of the production line encoder corresponding to the linear industrial camera and the area industrial camera is read. and And record the absolute physical coordinates of the production line encoder. ,in and The absolute physical coordinates of the production line encoder The corresponding values at the instant of acquisition of the two video feeds. Process parameter data from the production execution system are acquired synchronously to form a process feature set. The system synchronously reads the quality inspection records corresponding to the steel plate samples from the quality inspection system and selects historical high-quality sample sets according to the GB / T3274-2017 standard. And the linear array image sequence Area array image sequence Absolute displacement value of production line encoder and Absolute physical coordinates of production line encoders Process feature set and historical high-quality sample sets Each steel plate is bound to its corresponding sampling timestamp and then output. The process feature set is described below. This includes parameters such as the temperature of the heating furnace soaking zone, the tension fluctuation between stands, the rate of change of the rolling exit speed, and the water pressure distribution parameters of the laminar cooling zone. This represents the number of process parameters.
[0029] S1-2 Spatiotemporal sequence alignment and normalization processing. First, the linear array image sequence output from S1-1 is used. The array image sequence and the absolute displacement values of the production line recorded at the corresponding sampling time. and Input data and set the camera reference mounting distance to a fixed value of 500mm.
[0030] Secondly, calculate the physical spatial relative offset between linear and area array images. , It characterizes the true physical deviation between the two images in the observation area on the steel plate surface.
[0031] Subsequently, by the relative offset of the physical space Divide by the fixed physical size of the unit pixel, 0.1 mm, to calculate the pixel offset value. Linear image sequence Using the coordinates as a reference, a bicubic interpolation algorithm is employed, utilizing the pixel offset values. Opponent image sequence Resampling is performed to achieve sub-pixel spatial alignment, resulting in a spatially aligned area image sequence. ;Will Spatial aligned area array image sequence Scaling to uniform Pixel size and perform normalization.
[0032] Finally, the normalized linear image sequence and the area image sequence are concatenated using a feature stitching operation along the channel dimension, with the output dimension being... Aligned normalized feature map sequence .in, The linear array image sequence of the first channel is denoted as The second channel's array image sequence is denoted as .
[0033] S1-3 Normalized Surface State Feature Tensor Construction. This step uses the output of S1-1... and S1-2 output For input data.
[0034] First, based on Determine the spatial position of the current image frame on the production line, and calculate using a two-dimensional Gaussian distribution function. Irradiance value of each pixel within the sampling area The dimension is obtained as Ambient light intensity mapping map Its formula is: ; in, and The coordinate index represents the pixel's horizontal and vertical coordinates; the value 512 represents the coordinates of the image center point; the denominator value 524288 represents the variance constant of the corresponding light field diffusion feature. The irradiance weights are directly mapped to pixel-level values through a functional relationship, thereby reflecting the physical distribution characteristics of the light source within the imaging field of view, where the center is bright and the edges are dark.
[0035] Secondly, Mapping with ambient light intensity Perform feature concatenation operations according to the channel dimension to construct a dimension of... Standardized surface state feature tensor And output. The first channel stores linear array enhanced texture features. Geometric features of the second channel storage array Third channel storage environment light intensity mapping map .
[0036] S2. Construct a visual attention feature encoding network based on physical information constraints. This step aims to establish a feature encoding network that maps surface features of the steel plate to its intrinsic physical properties. Through deep coupling of visual attention mechanisms and physical conservation constraints, it achieves accurate extraction of defect features. The overall process is as follows: Figure 2 As shown: First, visual attention block segmentation and linear embedding are performed, and the standardized surface state feature tensor is then processed. Projection as block embedding matrix And superimpose the position encoding matrix Generate the initial input sequence Secondly, multi-head self-attention feature encoding is performed for the initial input sequence. Establish long-range feature dependencies and generate preliminary spatial feature maps using Transformer encoding blocks. ; Reconstruct the physical reflection component prediction branch, and use the initial spatial feature map Input into this branch and obtain the predicted intrinsic surface reflectance value. Subsequently, the physical information neural network PINN is introduced to predict the intrinsic reflectivity of the surface. Constraint correction is performed to reduce the irradiance of ambient light sources using physical imaging theory. The resulting impact on brightness gain, and the calculation of theoretically predicted intensity values. With image observation intensity Physical constraint loss function between Finally, the physical eigenvalue map is obtained based on physical consistency correction. and the preliminary spatial feature map Deep concatenation generates high-dimensional physical-visual joint feature vectors .
[0037] S2-1 Visual Attention Block Segmentation and Linear Embedding. This step uses the normalized surface state feature tensor output from S1-3. For input data. First, [the following is a list of input data]. Cut to Image blocks Each image block The size is Use a single fully connected layer for each image patch. The pixel values are flattened in channel order to obtain a one-dimensional vector of length 768. This vector is then linearly projected into a 512-dimensional embedding vector through a fully connected layer, resulting in a vector of dimension [missing value]. block embedding matrix The introduced dimension is... Learnable positional encoding matrix The Used to characterize the coordinate order of each image patch in physical space. By... and Perform element-wise addition to generate an initial input sequence carrying spatial location information. .
[0038] S2-2 Multi-head Self-Attention Feature Encoding. This step directly addresses the initial input sequence output by S2-1. Perform multi-level feature encoding. Construct a system containing 4... The feature extraction backbone, composed of concatenated coding blocks, is used to extract the initial input sequence. The input is fed into the first encoding block. Each encoding block contains a multi-head attention layer consisting of eight parallel self-attention heads. Each self-attention head computes a query vector. Key vector AND value vector The dot product weights are used to establish long-distance feature dependencies across image patches and across viewpoint channels. The sequence data after continuous mapping transformation through four coding blocks... According to the original The grid index is reshaped to restore the one-dimensional sequence to a two-dimensional matrix structure with a spatial resolution of . Preliminary spatial feature map with 512 channels .
[0039] S2-3 Physical Reflection Component Prediction Branch. This branch uses the preliminary spatial feature map output from S2-2 in this step. For input data. The aforementioned The input is fed into a parameter mapping network consisting of three consecutive two-dimensional convolutional layers. The first convolutional layer contains 32 kernels with a stride of 1; the second convolutional layer contains 64 kernels with a stride of 1, used for nonlinear fusion of high-dimensional semantic features; the third convolutional layer uses one kernel to perform channel compression. Through layer-by-layer linear and nonlinear mapping of the convolutional sequence, the 512-dimensional feature space is transformed into a single-channel intrinsic reflectance space, with an output dimension of... Predicted intrinsic surface reflectance .
[0040] S2-4 Physical Information Neural Network PINN Constraint Correction (1) Extracting the standardized surface state feature tensor Ambient light intensity mapping map stored in Real-time ambient light irradiance data were obtained by performing bilinear downsampling processing. ; This is used to provide the physical boundary conditions of the illumination at the current sampling coordinates. Next, the intensity values are predicted using physical imaging theory. Its calculation formula is defined as: ; in, The predicted intrinsic reflectance of the surface output by S2-3; The fixed camera observation angle is determined by the geometric parameters of the imaging unit. (2) For the standardized surface state feature tensor The first two channels recorded and Perform average fusion to obtain the image observation intensity. The image observation intensity The numerical unit is defined as a standardized brightness unit, whose value range is distributed in a continuous interval between 0 and 1, representing the true physical response value of the steel plate surface; (3) Calculate the theoretically predicted intensity value With the intensity of the image observation The sum of squared residuals between them is used to construct the physical constraint loss function. : ; in, The penalty coefficient is fixed at 0.5. Based on the aforementioned physical constraint loss function. The predicted intrinsic surface reflectance output from step S2-3 Perform physical consistency correction to reduce ambient light irradiance. The resulting brightness gain effect yields the physical eigenvalue map. .
[0041] S2-5 Feature Concatenation Output. This outputs a preliminary spatial feature map with long-range dependencies, derived from S2-2. With physical eigenvalue map Perform depthwise concatenation along the channel dimension to obtain a high-dimensional physical-visual joint feature vector. .
[0042] S3. Spatiotemporal Co-alignment and Correlation Modeling of Heterogeneous Defect and Process Data This step aims to establish a dynamic mapping relationship between defect visual features and production line process parameters, and to establish an evaluation baseline by quantifying the process contribution weight, providing strong constraint signals in the process dimension for subsequent topology reasoning. The overall process is as follows: First, perform spatiotemporal co-alignment of defect and process heterogeneous data, and use the absolute physical coordinates of the production line encoder obtained by S1. Perform an inverse mapping from the spatial domain to the temporal domain to obtain the defect feature vector. , defect feature vector Joint feature vectors with high-dimensional physical vision Perform splicing defect process associated state tensor Secondly, perform correlation reasoning based on a dual-track machine learning algorithm to analyze the defect process correlation state tensor. Process feature vectors are obtained by performing process feature channel extraction and average pooling operations. Using the regression coefficient matrix as Elastic network regression model and ensemble learning model Predict the continuous geometric index vector respectively With the defect category probability vector ; Perform process contribution quantification extraction again, combined with the regression coefficient matrix Compared with defect category probability vectors Attribution analysis to obtain the process correlation mapping matrix Finally, based on the historical high-quality sample set output from S1... Process feature set process feature vector and process correlation mapping matrix Calculate the historical optimal process configuration Incremental prediction results of defect traits Generate a baseline for evaluating process objectives. .
[0043] S3-1 Defect and Process Heterogeneous Data Spatiotemporal Co-alignment. This step uses the absolute physical coordinates of the production line encoder output in step S1-1. Process feature set and the high-dimensional physical-visual joint feature vector output from steps S2-5 For input data.
[0044] First, using the absolute physical coordinates of the production line encoder Perform an inverse mapping from the spatial domain to the temporal domain. The specific process of this inverse mapping is as follows: Let the physical coordinates corresponding to the current defect feature be... In the absolute physical coordinates of the production line encoder Chinese query and Encoder coordinates with minimum Euclidean distance According to the production line operating speed and the inherent sampling delay of the sensor A time window for process parameters strongly correlated with defect generation is determined. The time window is defined as an interval of... to( ;in for The corresponding timestamp, This is the time delay tolerance determined based on the thermal conductivity and plastic deformation characteristics of the steel plate.
[0045] Secondly, in the time window to( Internally, a mean statistical operation is performed on each process parameter to obtain the process feature vector corresponding to the current defect feature. Finally, the defect feature vector Joint feature vectors with high-dimensional physical vision Perform one-to-one pairing and splice along the channel dimension to generate a defect process-related state tensor. .
[0046] S3-2 Correlation Reasoning Based on a Dual-Track Machine Learning Algorithm. This step uses the defect process correlation state tensor output from step S3-1. For input data. First, from the defect process associated state tensor Extracting the process features corresponding to the post Each channel is used to perform average pooling along the spatial dimension to obtain the current process feature vector. Secondly, define the output attributes of the defect data: a vector of continuous geometric indices of defects. Includes vertical length, horizontal width, and prediction depth, all in mm; defect category probability vector. The dimension is 4, and its four components correspond to the predicted probabilities of four types of defects: cracks, folds, inclusions, and scratches.
[0047] (1) Elastic network regression inference for continuous defect geometric indices. The process feature vectors... Input an elastic network regression model and output the prediction results of a continuous geometric index. The regression coefficient matrix corresponding to the elastic network regression model is denoted as... .in, Columns 1, 2, and 3 represent the mapping relationship between each process parameter and the longitudinal length, transverse width, and predicted depth, respectively.
[0048] (2) Integrated classification reasoning based on defect category identification results. This involves integrating the process feature vectors... Input ensemble learning model Output defect probability vector The ensemble learning model It consists of 5 regression tree models, each with a maximum tree depth of 4 and a minimum number of split samples of 10.
[0049] S3-3 process contribution quantification extraction.
[0050] (1) First, the regression coefficient matrix The absolute values of each column element are taken, and an average operation is performed along the continuity geometric index dimension to obtain the process contribution weight vector corresponding to the continuity defect. .in, The The component represents the first... The combined linear contribution weight of each process parameter to the geometric indices of continuous defects.
[0051] (2) Using SHAP to analyze the probability vector of defect categories Attribution analysis is performed to obtain the process contribution weights corresponding to discrete defects. Specifically, this is done for the current process feature vector. The first in Each process parameter, integrated learning model Using the complete process feature vector as input, the predicted probability vectors of each type of defect are obtained through inference. ; then remove the first Each process parameter is again based on an ensemble learning model. Perform model inference to obtain the predicted probability vectors for each type of defect. Calculate the difference between the two predicted probability vectors to obtain the first... The contribution vector of each process parameter to the prediction of each type of defect in this sample. Regarding the contribution vector The four components are averaged to obtain the first... The scalar contribution of each process parameter to the prediction of the discrete defect category of this sample; perform this inference operation on all samples and take the average to obtain the first... Discrete contribution weights corresponding to each process parameter For all Repeat the above process for each process parameter to obtain the process contribution weight corresponding to discrete defects. .
[0052] (3) The process contribution weight vector corresponding to continuous defects Process contribution weight vector corresponding to discrete defects Perform element-wise averaging to obtain the overall process-related weight vector. Then, the overall process associated weight vector is used. Construct a diagonal process association mapping matrix The process association mapping matrix It is used to characterize the intensity of the effect of each process parameter deviating from the historical optimal process configuration on the fluctuation of defect characteristics.
[0053] S3-4 Calculates the baseline for evaluating process objectives. This step uses the historical high-quality sample set output from step S1-1. Process feature set The process feature vector extracted in step S3-2 and the process correlation mapping matrix output in step S3-3 For input data.
[0054] First, calculate the historical high-quality sample set. Each sample corresponds to a process feature set The historical optimal process configuration is obtained by statistically averaging the parameters of each dimension. .
[0055] Subsequently, the process feature vector is calculated. Compared to the historical optimal process configuration Deviation vector Next, matrix multiplication is performed. Obtain the prediction results of the defect trait increment, which reflects the change in defect severity. .
[0056] Finally, based on the aforementioned historical optimal process configuration Incremental prediction results of defect traits and predetermined fixed weighting coefficients Generate a baseline for evaluating process objectives. Among them, the weighting coefficient Optimization based on historical data: Collect production data for 30 consecutive days, with the optimization objective being to minimize the actual defect rate. A grid search method is used to traverse the interval [0,1] with a step size of 0.1. The selection process ultimately yields the result that minimizes the sum of the defect false positive rate and the defect missed detection rate on the validation set. The value is used as a fixed parameter of the system.
[0057] S4. Surface Defect Prediction of Steel Plates Based on Graph Attention Network (GAT) This step aims to establish a topological reasoning system based on graph neural networks, which integrates local visual texture and global process constraint information to achieve accurate prediction and category determination of defect areas; the overall process is as follows: Figure 3 As shown: First, a defect topology graph is constructed, and the high-dimensional physical-visual joint feature vector output by S2 is processed. The defect probability density map is obtained through upsampling and global average pooling operations. and visual node feature matrix Then, using the defect probability density map Candidate region nodes are extracted and an initial topology graph is constructed; then, process constraint feature mapping and joint node state construction are performed, and the process association mapping matrix output by S3 is applied. With process feature vector Perform a linear mapping to obtain the process risk characterization vector. The baseline for evaluating process objectives output by S3. Adding them together yields the process constraint description vector. Describing the process constraints as vectors With visual node feature matrix The visual-process joint node feature matrix is obtained by splicing. Finally, the initial topology graph and the vision-process joint node feature matrix are combined. The input graph attention network performs joint topological inference, and outputs a topology-enhanced feature matrix through two layers of graph attention propagation. Topology-enhanced feature matrix Input multi-task output head to obtain defect location mask Defect category identification results .
[0058] S4-1 Defect Topology Map Construction. This step uses the high-dimensional physical-visual joint feature vector output from S2-5. Using the input, construct a topology map of candidate defect regions.
[0059] First, Input to The probability prediction branch, consisting of convolutional layers and bilinear upsampling layers, with an upsampling factor of 16, yields a defect probability density map with the same resolution as the original input image. .in, Each pixel value represents the probability that the corresponding location belongs to a defect area.
[0060] Secondly, in Perform connected component analysis and extract all probabilities not less than a threshold. The locally connected region is defined as a defect candidate region. The threshold The value is fixed at 0.85, and each candidate region is treated as a node.
[0061] from Extract local features from the corresponding region: Visual node feature vectors are obtained after global average pooling. Stack the visual node feature vectors corresponding to all candidate regions row by row to obtain the visual node feature matrix: ; in, This represents the actual number of candidate regions within the current image frame.
[0062] For each candidate region Calculate its centroid coordinates in the original image pixel coordinate system. An initial adjacency matrix is constructed based on the spatial proximity relationship between the centroids of candidate regions. For any two candidate regions and Calculate the Euclidean distance between their centroids: ; when When pixels are used, let Otherwise, This leads to the construction of an initial topological map of the defect candidate regions: .in, This represents the set of candidate region nodes.
[0063] S4-2 Process Constraint Feature Mapping and Joint Node State Construction. This step uses the visual node feature matrix output from step S4-1. The current process feature vector extracted in step S3-2 The process correlation mapping matrix output in step S3-3 and the process target evaluation baseline output from steps S3-4. Using the input, construct the process constraint features corresponding to the candidate region nodes.
[0064] First, using the process association mapping matrix For the current process feature vector Perform a linear mapping to obtain the process risk representation vector corresponding to the current sample: ,in, .
[0065] Secondly, the process risk characterization vector Baseline of process target assessment Perform element-wise addition to obtain the process constraint description vector. ,in, .
[0066] Subsequently, the process constraint description vector copy This is repeated, and compared with the visual node feature matrix obtained in step S4-1. By concatenating the columns, we obtain the joint visual-process node feature matrix. ,in, Each row contains both the local visual features of the corresponding candidate region and the process constraint information of the current sample.
[0067] S4-3 Joint topology reasoning for vision and technology based on graph attention networks. This step uses the initial topology graph output from step S4-1. The joint visual and technological node feature matrix output from step S4-2 Using the input as input, a graph attention network inference module is constructed, outputting a topologically enhanced representation of the candidate regions. The graph attention network consists of two cascaded graph attention layers, each containing four independent attention heads, with each attention head having a fixed hidden feature dimension of 256. For nodes... Graph attention networks rely on adjacency matrices Determine its neighborhood node set, and perform weighted aggregation of the feature correlations between neighboring nodes to obtain the node. The updated feature representation is then processed through two layers of graph attention propagation, resulting in the output topology-enhanced feature matrix. The At the same time, the local texture information, spatial proximity information, and process constraint information of the candidate region are retained for subsequent multi-task output.
[0068] Based on this, a multi-task output head is constructed for... Performing joint inference yields the following output: First, split the branch pairs Perform step-by-step upsampling decoding and output a defect location mask. .in, In the middle, areas with a pixel value of 1 represent defective areas, and areas with a pixel value of 0 represent non-defective areas; Second, category decision branch pairs Perform global average pooling and fully connected mapping, and output the defect category identification result corresponding to the current defect entity. .
[0069] Therefore, the unified output of step S4-3 includes the defect location mask. , and trap category identification results The defect categories include four types: cracks, folds, inclusions, and scratches.
[0070] Experimental verification and analysis: To comprehensively verify the effectiveness and physical robustness of the proposed steel plate surface defect recognition method integrating visual attention mechanism, this embodiment constructs a simulation experimental platform that includes simultaneous acquisition of multi-source images, physical information-constrained feature encoding, and graph attention topological inference. The experimental dataset is constructed based on real historical monitoring data from a typical strip steel production line, and strictly follows the aforementioned standardized tensor construction logic, extracting 30,000 sets of standardized surface state feature tensors covering different lighting environments, complex rolling textures, and heterogeneous process parameter constraints. The experimental verification process focuses on evaluating the invention's overall accuracy in defect identification, its resistance to non-stationary illumination interference, and its threshold extraction based on candidate regions. The defects are extracted based on the three dimensions of reliability, which are used to quantify the performance.
[0071] Figure 4 The results of comparative tests on the performance of multiple algorithms in steel plate surface defect recognition are presented. The horizontal axis represents four different image processing and deep network model configurations, while the vertical axis represents the overall accuracy of defect recognition and the false alarm rate caused by non-stationary lighting. After processing by the architecture of this invention, the model achieves an overall recognition accuracy of 93.8% for four typical defects: cracks, folds, inclusions, and scratches, while reducing the false alarm rate to 4.2%. Using conventional digital image processing and basic convolutional neural networks, the accuracies are 78.5% and 85.2%, respectively, with false alarm rates both exceeding 10%. The above objective comparative data demonstrates that the deep coupling of the Physical Information Network (PINN) and the Graph Attention Network (GAT) can extract highly recognizable defect features and reduce the classification misclassification rate caused by non-stationary lighting and rolling textures in complex production line environments.
[0072] Figure 5The results of anti-reflective interference ablation experiments before and after the introduction of physical constraints are presented. In the figure, the horizontal axis represents the non-stationary reflective interference intensity index in the simulated production line environment, and the vertical axis represents the average cross-union ratio (CURRR) of defect pixel-level segmentation. Experimental results show that, using the baseline visual coding model without the physical reflection consistency equation, the average CURRR decreases rapidly as the reflective interference intensity increases, dropping to approximately 58% at an intensity of 80. After introducing the PINN constraint layer of the Physical Information Neural Network, the intensity value is predicted by forcing the network parameter updates to follow the physical imaging theory. With image observation intensity The model effectively isolates real-time ambient light irradiance data from the consistency patterns between them. The brightness gain achieved maintains an average cross-union ratio of 88% or higher under various reflective intensities. The test data above demonstrates that the physical constraint layer can assist the visual backbone network in correcting the nonlinear bias during photoelectric conversion, thereby improving recognition robustness.
[0073] Figure 6 Demonstrates threshold-based candidate region extraction The identification distribution correlation curve. The horizontal axis in the figure represents the defect probability density map. The output is a comprehensive confidence score, with the vertical axis representing the absolute pixel error in defect location prediction. When the comprehensive confidence score is below 0.85, the region is considered background noise or unstable interference; when the comprehensive confidence score is not less than 0.85, the absolute prediction error converges to an extremely low level, and the system successfully extracts candidate defect region nodes and performs subsequent topology inference. The above distribution data indicates that this invention, by setting a fixed threshold... With a value of 0.85, it can effectively intercept interference signals with uncertainties, ensuring the accuracy of the output defect location mask. With category recognition results It possesses a high degree of physical reliability.
[0074] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0075] While the above description illustrates specific embodiments of the present invention, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A method for identifying surface defects in steel plates by incorporating a visual attention mechanism, characterized in that, Includes the following processes: S1: Use linear and area array cameras to capture linear and area array image sequences of the steel plate surface, obtain the absolute physical coordinates of the production line encoder, perform spatiotemporal sequence alignment to obtain an aligned and normalized feature map sequence; use a two-dimensional Gaussian distribution function to establish an ambient light intensity mapping map and perform channel stitching to output a standardized surface state feature tensor with a defined dimension. Simultaneously acquire process feature sets and historical high-quality sample sets; S2, construct a visual attention feature encoding network that integrates physical information constraints, remove the influence of ambient light brightness gain on image quality, and obtain a high-dimensional physical-visual joint feature vector; S3, based on the absolute physical coordinates of the production line encoder and the joint feature vector of high-dimensional physical vision, performs defect and process heterogeneous data association modeling, quantifies the contribution weight of process parameters to defect formation through dual-track machine learning algorithm, obtains the current process feature vector and process association mapping matrix, and introduces process feature set and historical high-quality sample set to construct process target evaluation baseline. S4. Construct a steel plate surface defect detection network, and input the high-dimensional physical vision joint feature vector, the current process feature vector, the process correlation mapping matrix and the process target evaluation baseline into the network to obtain the defect detection results.
2. The method for identifying steel plate surface defects by incorporating a visual attention mechanism as described in claim 1, characterized in that: S1 specifically includes: First, multi-source sensor deployment and raw data acquisition are performed simultaneously. Linear array and area array cameras are used to capture linear array image sequences of the steel plate surface. With array image sequence And bind the absolute physical coordinates of the production line encoder. With process feature set Secondly, spatiotemporal sequence alignment and standardization are performed, utilizing the absolute displacement values of the production line. and Calculate the relative offset in physical space s and pixel offset value The spatial deviation between the two images is eliminated by subpixel-level interpolation resampling, and an aligned and normalized feature map sequence is output. Finally, a standardized surface state feature tensor is constructed, and an ambient light intensity mapping map is established using a two-dimensional Gaussian distribution function. It then performs channel splicing and outputs a standardized surface state feature tensor with defined dimensions. ; The process feature set This includes the temperature of the heating furnace soaking zone, the tension fluctuation value between stands, the rate of change of the rolling exit speed, and the water pressure distribution parameters of the laminar cooling zone. The number of process parameters; the historical high-quality sample set is formed by screening the quality inspection records corresponding to the steel plate samples in the quality inspection system during synchronous reading.
3. A method for identifying steel plate surface defects by incorporating a visual attention mechanism as described in claim 1 or 2, characterized in that: The specific process of performing spatiotemporal sequence alignment and normalization is as follows: Linear image sequence The array image sequence and the absolute displacement values of the production line recorded at the corresponding sampling time. and Input data and set the camera reference mounting distance to a fixed value of 500mm; Secondly, calculate the physical spatial relative offset between linear and area array images. ; Then, the physical space relative offset Divide by a fixed value of the physical size of a unit pixel to calculate the pixel offset value. Linear image sequence Using the coordinates as a reference, a bicubic interpolation algorithm is employed, utilizing the pixel offset values. Opponent image sequence Resampling is performed to obtain a spatially aligned area image sequence. ;Will Spatial aligned area array image sequence Perform uniform scaling and normalization; Finally, the normalized linear image sequence and the area image sequence are concatenated along the channel dimension to output an aligned and normalized feature map sequence. ;in, The linear array image sequence of the first channel is denoted as The second channel's array image sequence is denoted as .
4. The method for identifying steel plate surface defects by incorporating a visual attention mechanism as described in claim 3, characterized in that: The process of establishing an ambient light intensity mapping using a two-dimensional Gaussian distribution function and performing channel stitching to output a standardized surface state feature tensor with a defined dimension is as follows: Based on the absolute displacement value of the production line The spatial position of the current image frame on the production line is determined, and the irradiance value of each pixel within the sampling area is calculated using a two-dimensional Gaussian distribution function. The resulting image is a mapping map with the dimension of ambient light intensity. ; Align and normalize the feature map sequence Mapping with ambient light intensity Perform feature concatenation operations along the channel dimension to construct a standardized surface state feature tensor. And output; wherein, the first channel stores linear array enhanced texture features. Geometric features of the second channel storage array Third channel storage environment light intensity mapping map .
5. The method for identifying steel plate surface defects by incorporating a visual attention mechanism as described in claim 1, characterized in that: The visual attention feature encoding network in S2 is specifically as follows: First, visual attention block segmentation and linear embedding are performed to normalize the surface state feature tensor. Projection as block embedding matrix And superimpose the position encoding matrix Generate the initial input sequence Secondly, multi-head self-attention feature encoding is performed for the initial input sequence. Establish long-range feature dependencies and generate preliminary spatial feature maps using Transformer encoding blocks. ; Reconstruct the physical reflection component prediction branch and use the initial spatial feature map. Input into this branch and obtain the predicted intrinsic surface reflectance value. Subsequently, the physical information neural network PINN is introduced to predict the intrinsic reflectivity of the surface. Constraint correction is performed to reduce the irradiance of ambient light sources using physical imaging theory. The resulting impact on brightness gain, and the calculation of theoretically predicted intensity values. With image observation intensity Physical constraint loss function between Finally, the physical eigenvalue map is obtained based on physical consistency correction. and the preliminary spatial feature map Deep concatenation generates high-dimensional physical-visual joint feature vectors .
6. The method for identifying steel plate surface defects by incorporating a visual attention mechanism as described in claim 2, characterized in that: The specific process of S3 is as follows: First, perform spatiotemporal alignment of defect and process heterogeneous data, using the absolute physical coordinates of the production line encoder obtained from S1. Perform an inverse mapping from the spatial domain to the temporal domain to obtain the defect feature vector. , defect feature vector Joint feature vectors with high-dimensional physical vision Perform splicing defect process associated state tensor Secondly, perform correlation reasoning based on a dual-track machine learning algorithm to analyze the defect process correlation state tensor. Process feature vectors are obtained by performing process feature channel extraction and average pooling operations. Using the regression coefficient matrix as Elastic network regression model and ensemble learning model Predict the continuous geometric index vector respectively With the defect category probability vector ; Perform process contribution quantification extraction again, combined with the regression coefficient matrix Compared with defect category probability vectors Attribution analysis to obtain the process correlation mapping matrix Finally, based on the historical high-quality sample set output from S1... Process feature set process feature vector and process correlation mapping matrix Calculate the historical optimal process configuration Incremental prediction results of defect traits Generate a baseline for evaluating process objectives. .
7. The method for identifying steel plate surface defects by incorporating a visual attention mechanism as described in claim 6, characterized in that: The execution defects and spatiotemporal co-alignment with heterogeneous process data are specifically as follows: First, using the absolute physical coordinates of the production line encoder Perform the inverse mapping from the spatial domain to the temporal domain: Let the physical coordinates corresponding to the current defect feature be... In the absolute physical coordinates of the production line encoder Chinese query and Encoder coordinates with minimum Euclidean distance ; Based on production line operating speed and the inherent sampling delay of the sensor Determine the time window for process parameters that are strongly correlated with defect generation; the time window is defined as an interval of... to( ;in for The corresponding timestamp, The time delay tolerance is determined based on the thermal conductivity and plastic deformation characteristics of the steel plate; Secondly, within the time window to( Internally, a mean statistical operation is performed on each process parameter to obtain the process feature vector corresponding to the current defect feature. Finally, the defect feature vector Joint feature vectors with high-dimensional physical vision Perform one-to-one pairing and splice along the channel dimension to generate a defect process-related state tensor. .
8. The method for identifying steel plate surface defects by incorporating a visual attention mechanism as described in claim 7, characterized in that: The specific process of performing associative reasoning based on a dual-track machine learning algorithm is as follows: From the defect process associated state tensor Extracting the process features corresponding to the post Each channel is used to perform average pooling along the spatial dimension to obtain the current process feature vector. ; Secondly, define the output attributes of the defect data: a vector of continuous geometric indices of defects. Includes vertical length, horizontal width, and prediction depth; defect category probability vector. The dimension is 4, and its four components correspond to the predicted probabilities of four types of defects: cracks, folds, inclusions and scratches. Elastic network regression inference for continuous defect geometric indices: incorporating process feature vectors Input an elastic network regression model and output the prediction results of a continuous geometric index. The regression coefficient matrix corresponding to the elastic network regression model is denoted as... ;in, Columns 1, 2, and 3 represent the mapping relationship between each process parameter and the longitudinal length, transverse width, and predicted depth, respectively. Integrated classification reasoning based on defect category identification results: integrating process feature vectors Input ensemble learning model Output defect probability vector The ensemble learning model It consists of 5 regression tree models.
9. The method for identifying steel plate surface defects by incorporating a visual attention mechanism as described in claim 8, characterized in that: The specific process for extracting the contribution metric of the execution process is as follows: First, the regression coefficient matrix The absolute values of each column element are taken, and an average operation is performed along the continuity geometric index dimension to obtain the process contribution weight vector corresponding to the continuity defect. ;in, The The component represents the first... The combined linear contribution weight of each process parameter to the geometric indices of continuous defects; Secondly, SHAP is used to analyze the defect category probability vector. Attribution analysis is performed to obtain the process contribution weights corresponding to discrete defects; for the current process feature vector The first in Each process parameter, integrated learning model Using the complete process feature vector as input, the predicted probability vectors of each type of defect are obtained through inference. ; then remove the first Each process parameter is again based on an ensemble learning model. Perform model inference to obtain the predicted probability vectors for each type of defect. ; Calculate the difference between the two predicted probability vectors to obtain the first... The contribution vector of each process parameter to the prediction of each type of defect in this sample. ; For contribution vector The four components are averaged to obtain the first... The scalar contribution of each process parameter to the prediction of the discrete defect category of this sample; perform this inference operation on all samples and take the average to obtain the first... Discrete contribution weights corresponding to each process parameter ; for all Repeat the above process for each process parameter to obtain the process contribution weight corresponding to discrete defects. ; Finally, the process contribution weight vector corresponding to the continuity defect. Process contribution weight vector corresponding to discrete defects Perform element-wise averaging to obtain the overall process-related weight vector. Then, the overall process associated weight vector is used. Construct a diagonal process association mapping matrix .
10. The method for identifying steel plate surface defects by incorporating a visual attention mechanism as described in claim 2, characterized in that: S4 specifically includes: First, a defect topology graph is constructed, and the high-dimensional physical-visual joint feature vector output by S2 is processed. The defect probability density map is obtained through upsampling and global average pooling operations. and visual node feature matrix Then, using the defect probability density map Candidate region nodes are extracted and an initial topology graph is constructed; then, process constraint feature mapping and joint node state construction are performed, and the process association mapping matrix output by S3 is applied. With process feature vector Perform a linear mapping to obtain the process risk characterization vector. The baseline for evaluating process objectives output by S3. Adding them together yields the process constraint description vector. Describing the process constraints as vectors With visual node feature matrix The visual-process joint node feature matrix is obtained by splicing. Finally, the initial topology graph and the vision-process joint node feature matrix are combined. The input graph attention network performs joint topological inference, and outputs a topology-enhanced feature matrix through two layers of graph attention propagation. Topology-enhanced feature matrix Input multi-task output head to obtain defect location mask Defect category identification results The defect categories include four types: cracks, folds, inclusions, and scratches.