Metal material surface defect detection method and system based on multiple sensors
By employing a multi-sensor collaborative detection method, and utilizing graph convolutional feature fusion networks and feature hypergraph structures, the problem of insufficient exploration of multimodal data correlations in existing technologies is solved, achieving high precision and stability in the detection of surface defects in metallic materials.
Patent Information
- Application Number
- CN202511177228.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-11-28
AI Technical Summary
In existing technologies, methods for detecting surface defects in metallic materials rely on a single sensor or simple feature fusion, which makes it difficult to fully reflect the surface characteristics of the material and lacks in-depth exploration of the complex correlations between multimodal data, resulting in inaccurate and unstable detection results.
Optical sensors, ultrasonic sensors, and eddy current sensors are used to collect image data, acoustic feature data, and electromagnetic feature data of the surface of metal materials, respectively. A local correlation graph is constructed through a graph convolution feature fusion network. Adaptive weighted fusion is performed by combining the feature hypergraph structure and feature correlation tensor to achieve multi-level correlation modeling of multimodal data.
It significantly improves feature representation capabilities and detection accuracy, enhances the environmental adaptability and anti-interference ability of the detection system, and improves detection precision and generalization performance.
Smart Images

Figure CN121030664A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to defect detection technology, and more particularly to a method and system for detecting surface defects in metallic materials based on multiple sensors. Background Technology
[0002] With the increasing demands for product quality in industrial manufacturing, the accuracy and reliability of surface defect detection in metallic materials are becoming increasingly important. Existing detection methods primarily rely on single sensors to collect data, which struggles to comprehensively reflect the surface characteristics of materials; or they simply combine data from multiple sensors, lacking in-depth analysis of the complex correlations between multimodal data, resulting in inaccurate and unstable detection results. Furthermore, fluctuations in data quality and environmental interference during the detection process can also affect detection performance.
[0003] Currently used feature fusion methods often employ simple feature concatenation or weighted averaging strategies, failing to fully utilize the complementary information of multimodal data. Furthermore, existing methods generally suffer from insufficient feature extraction, limited ability to model intermodal interactions, and a lack of adaptability in the feature fusion process. These limitations make it difficult for detection systems to cope with complex and changing real-world conditions, affecting detection robustness. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a method and system for detecting surface defects in metallic materials based on multiple sensors, which can solve the problems in existing technologies.
[0005] A first aspect of the present invention provides a method for detecting surface defects in metallic materials based on multiple sensors, comprising: Image data, acoustic feature data, and electromagnetic feature data of the metal material surface are acquired by optical sensors, ultrasonic sensors, and eddy current sensors, respectively, and preprocessed accordingly. Preprocessed image data, acoustic feature data, and electromagnetic feature data are input into a graph convolutional feature fusion network. The three modal features are constructed into a local correlation graph through a multi-layer graph convolutional structure. In the local correlation graph, nodes represent the feature vectors of each modality, and edges represent the local dependencies between adjacent features. Local correlation feature data is obtained through cascaded graph convolution operations. A feature hypergraph structure is constructed using preprocessed image data, acoustic feature data, electromagnetic feature data, and the local correlation feature data. The feature hypergraph structure includes original modal feature nodes and locally correlated feature nodes, and the hyperedges of the feature hypergraph structure represent the combination relationship between nodes. A feature correlation tensor is constructed to capture global modal interaction patterns, and global correlation features between modal features are obtained through tensor decomposition. Based on an attention mechanism, the feature hypergraph structure and the global correlation features are adaptively weighted and fused to output the defect detection result.
[0006] Optionally, The steps for acquiring image data, acoustic feature data, and electromagnetic feature data of a metallic material surface using optical sensors, ultrasonic sensors, and eddy current sensors, respectively, include: The grayscale variance and signal-to-noise ratio of the image data, the spectral kurtosis and energy dispersion of the acoustic feature data, and the impedance change rate and phase fluctuation of the electromagnetic feature data are calculated and combined to obtain feature statistics. The feature statistics are converted into membership values using a Gaussian membership function, and the uncertainty score is calculated based on the membership values. The secondary sampling region is determined based on the spatial distribution characteristics of the uncertainty score. When the uncertainty score exceeds the preset threshold of the corresponding sensor, parameter adjustment is triggered, wherein: the imaging resolution of the optical sensor adopts an exponential adjustment function; the frequency bandwidth and pulse repetition frequency of the ultrasonic sensor adopt a linear adjustment function; the excitation frequency and scanning density of the eddy current sensor adopt a logarithmic adjustment function; the independent variable of the adjustment function is the uncertainty score, and the dependent variable is the ratio of the adjusted parameter to the initial parameter; the adjusted parameters are used to collect data in the secondary sampling area.
[0007] Optionally, The preprocessed image data, acoustic feature data, and electromagnetic feature data are input into a graph convolutional feature fusion network. A multi-layer graph convolutional structure is used to construct a local correlation graph of the three modal features. Nodes in the local correlation graph represent feature vectors of each modality, and edges represent local dependencies between adjacent features. The steps for obtaining local correlation feature data through cascaded graph convolution operations include: Feature extraction is performed on the preprocessed image data, acoustic feature data, and electromagnetic feature data to obtain image feature vectors, acoustic feature vectors, and electromagnetic feature vectors, respectively. The Gaussian similarity matrix between the image feature vectors, acoustic feature vectors, and electromagnetic feature vectors is calculated. An adjacency matrix is constructed based on the Gaussian similarity matrix, and the adjacency matrix is normalized to obtain a normalized adjacency matrix. A multi-layered cascaded graph convolutional network structure is constructed, comprising an intra-modal feature aggregation layer, a cross-modal feature interaction layer, and a global feature fusion layer. Graph convolution is performed by multiplying the normalized adjacency matrix with the input feature set and the learnable weight matrix. The intra-modal feature aggregation layer uses the normalized adjacency matrix to perform graph convolution on a single modality feature. The cross-modal feature interaction layer uses the normalized adjacency matrix to interact and fuse features from different modalities. The global feature fusion layer uses the normalized adjacency matrix to integrate multi-modal features into a unified representation. A channel attention mechanism is introduced into the graph convolutional network structure. The importance score of the feature channel is calculated through a learnable weight matrix, and the feature is weighted and enhanced according to the importance score of the feature channel. The output features of the graph convolutional network structure are subjected to feature fusion processing to obtain local correlation feature data.
[0008] Optionally, The steps for constructing a multi-layered cascaded graph convolutional network structure include: Image feature vectors, acoustic feature vectors, and electromagnetic feature vectors are input into the intra-modal feature aggregation layer. The intra-modal weight vector is calculated by a multilayer perceptron. The image feature vectors, acoustic feature vectors, and electromagnetic feature vectors are weighted and aggregated based on the intra-modal weight vectors and the normalized adjacency matrix to obtain the intra-modal features. Intermodal attention maps are obtained by calculating normalized inner products between different modal features. Based on the intermodal attention maps, the intramodal features are mapped to value features. The value features are then weighted and combined to obtain cross-modal interaction features. The cross-modal interaction features are pooled at multiple scales to obtain a multi-scale feature set. The dynamic weights of each scale feature in the multi-scale feature set are calculated. The cross-modal interaction features are added to the multi-scale feature set weighted by the dynamic weights to obtain a global fusion feature. Construct inter-layer skip connections, concatenate the output features of the previous layer with the features of the current layer to calculate the gating coefficient, and dynamically fuse the output features of the previous layer and the features of the current layer based on the gating coefficient to obtain updated features; Calculate the mutual information between the features of each layer and the target features, determine the optimal number of network layers based on the mutual information, and adaptively weight the features of each layer to obtain the final fused features.
[0009] Optionally, The step of constructing a feature hypergraph structure using preprocessed image data, acoustic feature data, electromagnetic feature data, and the locally associated feature data, wherein the feature hypergraph structure includes original modal feature nodes and locally associated feature nodes, and the hyperedges of the feature hypergraph structure represent the combination relationships between nodes includes: A node set is constructed using preprocessed image data, acoustic feature data, electromagnetic feature data, and local correlation feature data. An initial hyperedge set is generated based on the node set. The hyperedge weights are initialized using a multilayer perceptron. The initial node set is then weighted and aggregated to obtain the initial node representation. A mutual information matrix is constructed based on the initial node representation. A higher-order correlation tensor is constructed using the mutual information matrix. The node features are then enhanced based on the higher-order correlation tensor to obtain enhanced node features. Multi-scale pooling is performed on the enhanced node features to obtain a multi-layer feature pyramid. A feature mapping matrix between adjacent layers is constructed, and weighted fusion is performed based on the feature mapping matrix to obtain the fused feature representation. Cluster center nodes are determined by calculating node density and relative distance based on fused feature representation, a dynamic hyperedge structure is constructed, and the adaptive weights of the dynamic hyperedge structure are calculated. A temporal state transition function is constructed, and the node state and hyperedge structure are updated temporally based on the temporal state transition function. The updated nodes are mapped to knowledge graph entities to obtain knowledge features, and the knowledge features are fused with the node features to obtain an enhanced feature representation. A joint optimization objective is constructed based on reconstruction loss, structure preservation loss, and temporal consistency loss. The joint optimization objective is then optimized to obtain the final dynamic adaptive hypergraph structure. The enhanced feature representation is then input into the dynamic adaptive hypergraph structure for feature propagation.
[0010] Optionally, The steps of calculating node density and relative distance based on fused feature representation to determine cluster center nodes, constructing dynamic hyperedge structures, and calculating the adaptive weights of the dynamic hyperedge structures include: Calculate the local density value and relative distance value between nodes, determine the node score based on the product of the local density value and the relative distance value between nodes, and determine the node center node if the node score is greater than the preset score threshold; based on the initial cluster center node, calculate the structural similarity and feature similarity of the node pairs, and obtain the comprehensive similarity by weighting the structural similarity and feature similarity. Based on the comprehensive similarity, nodes are grouped to obtain candidate hyperedges, the boundary nodes of the candidate hyperedges are identified, and the boundary nodes are added to the candidate hyperedges with the highest similarity to obtain optimized candidate hyperedges. An initial similarity threshold is determined based on the distribution of the comprehensive similarity. The initial similarity threshold is then adaptively adjusted in conjunction with the temporal stability of the hyperedge structure to obtain a final similarity threshold. The optimized candidate hyperedges are then selected based on the final similarity threshold. The degree of structural change is determined based on the node state changes and the hyperedge stability assessment results. Based on the degree of structural change, a corresponding update strategy is selected to dynamically adjust the hyperedge structure. For the adjusted hyperedge structure, the structural weight, feature weight, and temporal weight of the hyperedge are calculated. The final adaptive weight of the hyperedge is obtained by joint optimization based on local consistency, global structure preservation, and temporal smoothness.
[0011] Optionally, The steps for constructing a feature association tensor to capture global modal interaction patterns, and obtaining the global association features between modal features through tensor decomposition, include: The image feature matrix, acoustic feature matrix, and electromagnetic feature matrix are constructed into a third-order feature correlation tensor. The third-order feature correlation tensor is then standardized and sparsified to obtain a preprocessed tensor. The video-sound, video-electronic, and sound-electronic interaction matrices in the preprocessed tensor are calculated by tensor product operation. The norms of the three interaction matrices are calculated to obtain the local association strength. The local association strengths are aggregated to obtain the global association pattern. The preprocessed tensor is optimized based on the global association pattern. The optimized tensor is subjected to singular value decomposition. The optimal decomposition rank is determined based on the energy retention rate and a preset energy threshold. The optimal decomposition rank is used to decompose the tensor to obtain a group of factor vectors. The importance weight of the pattern is calculated based on the norm of the factor vectors. The group of factor vectors is then weighted and combined to obtain the global association feature. The norm of the global association feature is calculated to obtain the regularization constraint, and the difference norm between the global association feature and the local association strength is calculated to obtain the consistency constraint. The regularization constraint and the consistency constraint are weighted and combined to construct the feature optimization objective, and the optimized global association feature is obtained. The feature difference norm is calculated between the current global correlation feature and the global correlation feature at the previous time step to obtain the feature difference degree. When the feature difference degree is greater than the adaptive update threshold, the current global correlation feature and the global correlation feature at the previous time step are weighted and fused to update the global correlation feature.
[0012] Secondly, a multi-sensor-based surface defect detection system for metallic materials is provided, including: The first unit is used to acquire image data, acoustic feature data and electromagnetic feature data of the surface of the metal material through optical sensors, ultrasonic sensors and eddy current sensors respectively, and to perform preprocessing on each data. The second unit is used to input the preprocessed image data, acoustic feature data and electromagnetic feature data into the graph convolution feature fusion network. The three modal features are constructed into a local correlation graph through a multi-layer graph convolution structure. The nodes in the local correlation graph represent the feature vectors of each modality, and the edges represent the local dependencies between adjacent features. The local correlation feature data is obtained through cascaded graph convolution operations. The third unit is used to construct a feature hypergraph structure using preprocessed image data, acoustic feature data, electromagnetic feature data, and the local correlation feature data. The feature hypergraph structure includes original modal feature nodes and local correlation feature nodes, and the hyperedges of the feature hypergraph structure represent the combination relationship between nodes. A feature correlation tensor is constructed to capture the global modal interaction pattern, and the global correlation features between modal features are obtained through tensor decomposition. The feature hypergraph structure and the global correlation features are adaptively weighted and fused based on an attention mechanism to output the defect detection result.
[0013] Thirdly, a computer-readable storage medium is provided, having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0014] This invention achieves multi-level correlation modeling of multimodal data by constructing a graph convolutional feature fusion network and a feature hypergraph structure, effectively capturing local dependencies and combination patterns. Simultaneously, it introduces feature correlation tensors and tensor decomposition methods to deeply mine global modal interaction features. It innovatively performs adaptive fusion of local structural features and global correlation features, significantly improving feature representation capabilities and detection accuracy.
[0015] This invention employs a multi-sensor collaborative detection strategy to overcome the limitation of information from a single sensor. Through adaptive data acquisition and dynamic feature fusion mechanisms, the environmental adaptability and anti-interference capability of the detection system are enhanced. This method ensures detection accuracy while exhibiting good generalization performance, and can be widely applied to various surface defect detection scenarios for metallic materials. Attached Figure Description
[0016] Figure 1 This is a schematic flowchart of a multi-sensor-based method for detecting surface defects in metallic materials according to an embodiment of the present invention. Figure 2 Flowchart of dynamic hyperedge structure construction and adaptive weight calculation. Detailed Implementation
[0017] The technical solutions of the present invention will be described below with reference to the accompanying drawings. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0018] Figure 1 This is a schematic flowchart of the multi-sensor-based method for detecting surface defects in metallic materials according to the present invention. Figure 1 As shown, the method includes: Image data, acoustic feature data, and electromagnetic feature data of the metal material surface are acquired by optical sensors, ultrasonic sensors, and eddy current sensors, respectively, and preprocessed accordingly. The preprocessed image data, acoustic feature data, and electromagnetic feature data are input into a graph convolutional feature fusion network. The graph convolutional feature fusion network constructs a local correlation graph of the three modal features through a multi-layer graph convolutional structure. In the local correlation graph, nodes represent the feature vectors of each modality, and edges represent the local dependencies between adjacent features. Local correlation feature data is obtained through cascaded graph convolution operations. A feature hypergraph structure is constructed using preprocessed image data, acoustic feature data, electromagnetic feature data, and the local correlation feature data. The feature hypergraph structure includes original modal feature nodes and locally correlated feature nodes, and the hyperedges of the feature hypergraph structure represent the combination relationship between nodes. A feature correlation tensor is constructed to capture global modal interaction patterns, and global correlation features between modal features are obtained through tensor decomposition. Based on an attention mechanism, the feature hypergraph structure and the global correlation features are adaptively weighted and fused to output the defect detection result.
[0019] Optionally, The steps for acquiring image data, acoustic feature data, and electromagnetic feature data of a metallic material surface using optical sensors, ultrasonic sensors, and eddy current sensors, respectively, include: The grayscale variance and signal-to-noise ratio of the image data, the spectral kurtosis and energy dispersion of the acoustic feature data, and the impedance change rate and phase fluctuation of the electromagnetic feature data are calculated and combined to obtain feature statistics. The feature statistics are converted into membership values using a Gaussian membership function. An uncertainty score is calculated based on the membership values, and the uncertainty score characterizes the reliability of the detection data. A secondary sampling region is determined based on the spatial distribution characteristics of the uncertainty score, and the secondary sampling region is the region where the uncertainty score is higher than a preset score threshold. When the uncertainty score exceeds the preset threshold of the corresponding sensor, parameter adjustment is triggered, wherein: the imaging resolution of the optical sensor adopts an exponential adjustment function; the frequency bandwidth and pulse repetition frequency of the ultrasonic sensor adopt a linear adjustment function; and the excitation frequency and scanning density of the eddy current sensor adopt a logarithmic adjustment function; the independent variable of the adjustment function is the uncertainty score, and the dependent variable is the ratio of the adjusted parameter to the initial parameter; The adjusted parameters are used to acquire data from the secondary sampling area, obtaining high-resolution image data from the optical sensor, high-sampling-frequency acoustic feature data from the ultrasonic sensor, and high-density scanning electromagnetic feature data from the eddy current sensor, and then preprocessing them respectively.
[0020] For example, the optical sensor employs a high-resolution industrial camera equipped with a ring LED light source. Its operating resolution is dynamically adjustable from 1280×960 to 4096×3072 pixels, with an initial setting of 2048×1536 pixels. The light source brightness is 5000 lumens, and the exposure time is 20 milliseconds. The ultrasonic sensor uses a piezoelectric crystal transducer with an initial operating frequency of 5MHz, a frequency bandwidth of 2MHz, a pulse repetition frequency of 1000Hz, a sound wave incident angle of 45 degrees, and a sampling rate of 100MHz. The eddy current sensor uses a flat coil structure with an initial excitation frequency of 250kHz, 200 coil turns, a scan density of 10mm / step, a sampling rate of 10MHz, and a gain of 45dB.
[0021] Gray-level variance is obtained by calculating the standard deviation of the gray-level values of image pixels. Signal-to-noise ratio (SNR) is calculated as the ratio of effective signal energy to noise energy, generally requiring a value greater than 25 dB. The quality of acoustic feature data is assessed through spectral kurtosis and energy dispersion. Spectral kurtosis reflects the sharpness of the spectral distribution and is calculated as the ratio of the fourth moment to the square of the second moment, with a normal value between 3.0 and 5.0. Energy dispersion measures the degree of dispersion of sound wave energy in the frequency domain and is calculated as the standard deviation of the spectral energy distribution, with a typical value between 0.3 and 0.7. The quality of electromagnetic feature data is assessed through impedance change rate and phase fluctuation. Impedance change rate calculates the percentage change in impedance values between adjacent scan points, with a normal range of 0.5% to 2%. Phase fluctuation calculates the standard deviation of the phase angle, generally requiring a value less than 3 degrees. These six indicators are combined to form a feature statistics vector, which is then converted into membership values using a Gaussian membership function. The center point of the Gaussian membership function is set as the ideal value for each indicator. For example, the ideal value for image signal-to-noise ratio is 35 dB, with a standard deviation of 5 dB; the ideal value for spectral kurtosis is 4.0, with a standard deviation of 0.5; and the ideal value for impedance change rate is 1%, with a standard deviation of 0.3%. The membership value calculated by the Gaussian function is between 0 and 1, with values closer to 1 indicating higher feature quality. For each sensor, the uncertainty score is calculated by combining the membership values of its corresponding indicators, calculated by subtracting the average of the membership values from 1. The uncertainty score ranges from 0 to 1, with higher values indicating higher data uncertainty and lower reliability.
[0022] The sampling area was divided into a 5mm × 5mm grid, and a heatmap of the spatial distribution of uncertainty scores was plotted. A region growing algorithm was used to determine the secondary sampling region, i.e., the connected regions with uncertainty scores higher than a preset threshold. The preset score threshold was set according to the required detection accuracy: 0.3 for high-precision detection, 0.5 for medium-precision detection, and 0.7 for general-precision detection. In the experiment, during the detection of an aluminum alloy sheet, the average uncertainty score of the upper right corner region was found to be 0.62, exceeding the preset threshold of 0.5. This region was determined to be a secondary sampling region, with an area of approximately 120 square centimeters.
[0023] When the uncertainty score exceeds the sensor's preset threshold, adaptive parameter adjustment is triggered. The preset threshold for the optical sensor is 0.4. When the uncertainty score reaches 0.58, the imaging resolution is adjusted using an exponential adjustment function. The exponential adjustment function is expressed as: the ratio of the adjusted resolution to the initial resolution is equal to a power of 2, where the exponent is twice the uncertainty score minus the threshold. For example, if the uncertainty score is 0.58, exceeding the threshold by 0.18, the calculated adjustment coefficient is 2 to the power of 0.36, approximately 1.28. The initial resolution is 2048×1536 pixels, which is increased to 2621×1966 pixels after adjustment. The preset threshold for the ultrasonic sensor is 0.45. When the uncertainty score is 0.67, the frequency bandwidth and pulse repetition frequency are updated using a linear adjustment function. The linear adjustment function is expressed as: the adjustment coefficient is equal to 1 plus three times the uncertainty score minus the threshold. The calculated adjustment factor was 1.66, the initial frequency bandwidth was adjusted from 2MHz to 3.32MHz, and the pulse repetition frequency increased from 1000Hz to 1660Hz. The preset threshold for the eddy current sensor was 0.5. When the uncertainty score was 0.75, the excitation frequency and scan density were optimized using a logarithmic adjustment function. The logarithmic adjustment function is expressed as: the adjustment factor equals 1 plus the logarithm of the uncertainty score minus the threshold multiplied by 2. The calculated adjustment factor was 1.46, the initial excitation frequency was increased from 250kHz to 365kHz, and the scan density increased from 10mm / step to 6.85mm / step.
[0024] Data was acquired from the secondary sampling area using adjusted parameters. The optical sensor acquired images at a higher resolution of 2621×1966 pixels, increased the exposure time to 25 milliseconds, and increased the light source brightness to 6000 lumens to ensure detail clarity. The ultrasonic sensor scanned using a wider frequency bandwidth of 3.32MHz and a higher pulse repetition frequency of 1660Hz, increasing the sampling rate to 150MHz for more precise defect identification. The eddy current sensor detected with a higher excitation frequency of 365kHz and a denser scanning spacing of 6.85mm / step, increasing the sampling rate to 15MHz and adjusting the gain to 50dB to improve sensitivity to small defects.
[0025] After acquisition, the acquired high-quality data undergoes preprocessing. Image data preprocessing includes adaptive histogram equalization to enhance contrast, Gaussian filtering to remove noise, edge-preserving filtering to retain details, and normalization to the 0-1 range. Acoustic feature data preprocessing includes bandpass filtering to remove low-frequency drift and high-frequency noise, envelope detection to extract signal features, time-domain synchronous averaging to improve the signal-to-noise ratio, and wavelet denoising to retain effective information. Electromagnetic feature data preprocessing includes baseline correction to eliminate system drift, phase calibration to ensure measurement consistency, median filtering to remove outliers, and background suppression to highlight defect signals.
[0026] The double sampling strategy for uncertainty assessment in this invention can intelligently identify areas with poor data quality and perform targeted optimization, significantly improving the ability to identify different types of defects. The differentiated parameter adjustment function design allows each sensor to be optimally configured according to its own characteristics, achieving efficient utilization of detection resources and a significant improvement in detection accuracy, providing reliable technical support for defect detection in complex industrial environments.
[0027] Optionally, The preprocessed image data, acoustic feature data, and electromagnetic feature data are input into a graph convolutional feature fusion network. This network constructs a local correlation graph from the three modal features using a multi-layer graph convolutional structure. Nodes in the local correlation graph represent feature vectors of each modality, and edges represent local dependencies between adjacent features. The steps for obtaining locally correlated feature data through cascaded graph convolution operations include: Feature extraction is performed on the preprocessed image data, acoustic feature data, and electromagnetic feature data to obtain image feature vectors, acoustic feature vectors, and electromagnetic feature vectors, respectively. The Gaussian similarity matrix between the image feature vectors, acoustic feature vectors, and electromagnetic feature vectors is calculated. An adjacency matrix is constructed based on the Gaussian similarity matrix, and the adjacency matrix is normalized to obtain a normalized adjacency matrix. A multi-layered cascaded graph convolutional network structure is constructed, comprising an intra-modal feature aggregation layer, a cross-modal feature interaction layer, and a global feature fusion layer. Graph convolution is performed by multiplying the normalized adjacency matrix with the input feature set and a learnable weight matrix. The input feature set consists of image feature vectors, acoustic feature vectors, and electromagnetic feature vectors. The intra-modal feature aggregation layer uses the normalized adjacency matrix to perform graph convolution on single-modal features. The cross-modal feature interaction layer uses the normalized adjacency matrix to interact and fuse features from different modalities. The global feature fusion layer uses the normalized adjacency matrix to integrate multi-modal features into a unified representation. A channel attention mechanism is introduced into the graph convolutional network structure. The importance score of the feature channel is calculated through a learnable weight matrix, and the feature is weighted and enhanced according to the importance score of the feature channel. The output features of the graph convolutional network structure are subjected to feature fusion processing to obtain local correlation feature data.
[0028] For example, the preprocessed image data is a 256×256 pixel grayscale image, the acoustic feature data is a 5000-point time-domain signal, and the electromagnetic feature data is a 2000-point impedance and phase sequence. Feature extraction is performed on these three types of data separately. Image feature extraction uses an improved convolutional neural network containing five convolutional layers with 32, 64, 128, 256, and 512 kernels per layer, a 3×3 kernel size, and a stride of 1. Each layer is followed by a max-pooling layer with a 2×2 pooling size and a stride of 2. Finally, a 128-dimensional image feature vector is output through a fully connected layer. Acoustic feature extraction uses a joint time-frequency domain analysis method. A short-time Fourier transform is performed on the time-domain signal to generate a spectrogram with a size of 128×64. Then, a three-layer convolutional network is used to extract spectral features. The parameters of the convolutional layers are similar to those in image processing, but the number of channels is halved. Finally, a 64-dimensional acoustic feature vector is output. Electromagnetic feature extraction employs a recurrent neural network structure, including a bidirectional long short-term memory network with 128 hidden units, a time step of 50, and a fully connected output layer, generating a 96-dimensional electromagnetic feature vector.
[0029] Gaussian similarity matrices are calculated for the three extracted feature vectors to construct an intermodal association structure. The calculation method is as follows: for any two feature vectors, their Euclidean distance is calculated, and then converted into a similarity value using a Gaussian kernel function. The kernel width parameter of the Gaussian kernel function is set to 0.5 times the average distance between the feature vectors, with an empirical value of approximately 5.2. For example, the Gaussian similarity calculated between the image feature vector and the acoustic feature vector of a steel surface inspection sample is 0.73, the Gaussian similarity between the image feature vector and the electromagnetic feature vector is 0.65, and the Gaussian similarity between the acoustic feature vector and the electromagnetic feature vector is 0.81, indicating strong complementarity between acoustic and electromagnetic features. Based on the calculated Gaussian similarity matrix, a similarity threshold of 0.6 is set, and feature pairs with similarity higher than the threshold are connected to form an adjacency matrix. The size of the adjacency matrix is the square of the total number of feature nodes. For the case with 3 sample points, each sample point has 3 modal features, and the adjacency matrix size is 9×9. To maintain the numerical stability of graph convolution operations, the adjacency matrix is normalized. The normalization method is as follows: calculate the degree (number of connections) of each node, construct a degree matrix (a diagonal matrix, where the diagonal elements are the degrees of the nodes), multiply the adjacency matrix by the left power of the degree matrix to the power of -0.5, and then multiply it by the right power of the degree matrix to obtain the normalized adjacency matrix.
[0030] A multi-layered cascaded graph convolutional network structure is constructed based on the normalized adjacency matrix. The network contains three functional layers: an intra-modal feature aggregation layer, a cross-modal feature interaction layer, and a global feature fusion layer. The intra-modal feature aggregation layer processes single-modal features using graph convolution operations, preserving the integrity of intra-modal information. Specifically, the input feature matrix is multiplied by the normalized adjacency matrix, and then multiplied by the learnable weight matrix. The weight matrix dimensions for image features are 128×64, for acoustic features it is 64×32, and for electromagnetic features it is 96×48. The activation function of the intra-modal feature aggregation layer is LeakyReLU, with a negative half-axis slope of 0.2. The cross-modal feature interaction layer is designed to capture the relationships between features from different modalities. It is implemented by constructing a cross-modal adjacency submatrix, retaining only the connections between nodes from different modalities, and then performing graph convolution operations similar to those used for intra-modal aggregation. The weight matrix of this layer has a dimension of (64+32+48)×96, and the activation function is the Sigmoid function. The global feature fusion layer integrates the features output from the first two layers to generate a unified feature representation. This layer uses a residual connection structure, adding the original features to the features after graph convolution. The weight matrix has a dimension of 96×128, and the activation function is ReLU.
[0031] A channel attention mechanism is introduced into the graph convolutional network structure to enhance the representation ability of important feature channels. The channel attention mechanism consists of three steps: feature aggregation, importance evaluation, and feature reweighting. Feature aggregation processes the feature map using global average pooling and global max pooling respectively, generating two channel descriptors. The dimension of each descriptor is equal to the number of feature channels. Taking the 128-channel feature output from the global fusion layer as an example, two 128-dimensional channel descriptors are generated. Importance evaluation is implemented through a two-layer perceptron with shared weights. The first layer reduces the number of channels to 1 / 16 of the original (i.e., 8 neurons) and uses ReLU activation; the second layer restores the original number of channels (128 neurons). After the two channel descriptors are processed by the same perceptron, two sets of channel importance scores are obtained. These scores are summed and normalized to the 0-1 range using the sigmoid function to obtain the final channel importance weights. Feature reweighting multiplies each channel of the original feature by its corresponding importance weight, achieving selective enhancement of feature channels. For example, for crack defect samples, the edge feature channel weight of the visual modality reaches 0.92, while the texture feature channel weight is only 0.37; for pit defect samples, the depth feature channel weight of the ultrasonic modality reaches 0.88, indicating that the attention mechanism can adaptively emphasize important features according to different defect types.
[0032] The output features of the graph convolutional network structure are fused to generate the final locally correlated feature data. The fusion process includes three stages: feature concatenation, dimensionality adjustment, and normalization. Feature concatenation connects intra-modal features, cross-modal features, and global features along the channel dimension, resulting in a high-dimensional feature vector with dimensions (64+96+128)=288. Dimensionality adjustment maps the high-dimensional features to a unified 160-dimensional space through fully connected layers, reducing redundancy and controlling model complexity. Normalization uses batch normalization, calculating the mean and variance of the feature vector along the batch dimension. The mean is subtracted from the feature vector, divided by the square root of the variance, multiplied by a learnable scaling parameter, and a learnable bias parameter is added. The initial value of the scaling parameter is set to 1, and the initial value of the bias parameter is set to 0. After the above processing, the final locally correlated feature data is obtained, which contains the local dependencies of multimodal information, providing a foundation for subsequent hypergraph construction and feature fusion. For sample processing with a batch size of 32, the shape of the local correlation feature data is 32×160, stored as a floating-point array, and the value range is between -1 and 1.
[0033] This invention effectively captures local dependencies between multimodal features by constructing local correlation graphs and multi-layer graph convolutional structures, achieving the organic fusion of heterogeneous information. The introduction of a channel attention mechanism enables the network to adaptively emphasize important feature channels, improving the discriminative power of feature representations. The design of the multi-layer cascaded structure ensures the integrity of information within modalities and the sufficiency of interactions between modalities.
[0034] Optionally, The steps for constructing a multi-layered cascaded graph convolutional network structure include: Image feature vectors, acoustic feature vectors, and electromagnetic feature vectors are input into the intra-modal feature aggregation layer. The intra-modal weight vector is calculated by a multilayer perceptron. The image feature vectors, acoustic feature vectors, and electromagnetic feature vectors are weighted and aggregated based on the intra-modal weight vectors and the normalized adjacency matrix to obtain the intra-modal features. Intermodal attention maps are obtained by calculating normalized inner products between different modal features. Based on the intermodal attention maps, the intramodal features are mapped to value features. The value features are then weighted and combined to obtain cross-modal interaction features. The cross-modal interaction features are pooled at multiple scales to obtain a multi-scale feature set. The dynamic weights of each scale feature in the multi-scale feature set are calculated. The cross-modal interaction features are added to the multi-scale feature set weighted by the dynamic weights to obtain a global fusion feature. Construct inter-layer skip connections, concatenate the output features of the previous layer with the features of the current layer to calculate the gating coefficient, and dynamically fuse the output features of the previous layer and the features of the current layer based on the gating coefficient to obtain updated features; Calculate the mutual information between the features of each layer and the target features, determine the optimal number of network layers based on the mutual information, and adaptively weight the features of each layer to obtain the final fused features.
[0035] For example, image feature vectors, acoustic feature vectors, and electromagnetic feature vectors are input into the intra-modal feature aggregation layer. This layer calculates the intra-modal weight vector using a multilayer perceptron. The perceptron contains two hidden layers: the first layer has half the number of neurons in each feature dimension, and the second layer has one-quarter the number of neurons in each feature dimension. The image feature vector has a dimension of 128, corresponding to a multilayer perceptron structure of 128-64-32-128; the acoustic feature vector has a dimension of 64, corresponding to a structure of 64-32-16-64; and the electromagnetic feature vector has a dimension of 96, corresponding to a structure of 96-48-24-96. The weights are initialized using a truncated normal distribution with a mean of 0 and a standard deviation of 0.02, and the activation function is ReLU. Taking a steel plate crack sample as an example, the calculated weight vector values within the image feature mode range from 0.35 to 0.88, the weight vector within the acoustic feature mode ranges from 0.29 to 0.75, and the weight vector within the electromagnetic feature mode ranges from 0.42 to 0.91. Based on these weight vectors and a normalized adjacency matrix, the original feature vectors are weighted and aggregated to generate intra-modal features. The weighting method involves element-wise multiplication of the weight vector and the feature vector, followed by message passing through the normalized adjacency matrix. The resulting intra-modal features have the same dimension as the original feature vectors, but have undergone local structure awareness enhancement, incorporating information from neighboring nodes.
[0036] An intermodal interaction layer is constructed, and an intermodal attention map is obtained by calculating the normalized inner product between features of different modalities. The normalized inner product is calculated by performing a dot product operation on two feature vectors and then dividing by the product of the norms of the two vectors. Taking a batch size of 16 as an example, the generated intermodal attention map is 16×3×3 in size, where 3 represents the number of the three modes, and the diagonal elements of each 3×3 submatrix are 1 (representing self-attention), while the off-diagonal elements represent the attention weights between different modes. For example, in the intermodal attention map calculated from a pit defect sample on the surface of a metal plate, the image-acoustic attention weight is 0.72, the image-electromagnetic attention weight is 0.65, and the acoustic-electromagnetic attention weight is 0.81, indicating that the acoustic and electromagnetic modes have strong complementarity. Based on the intermodal attention map, intramodal features are mapped to value features. The mapping process is as follows: A query matrix, key matrix, and value matrix are constructed. The query matrix and key matrix are multiplied by a dot product to calculate the attention score. The attention score is then normalized using a softmax function and weighted and summed with the value matrix to obtain the value features. The query matrix, key matrix, and value matrix are all implemented using a linear transformation layer, with the dimensions remaining unchanged before and after the transformation. Taking image features as an example, the weight matrix of the linear transformation layer has a dimension of 128×128. The generated value features are then weighted and combined using a multi-head connection mechanism to obtain cross-modal interaction features. The number of heads is set to 4, and the feature dimension of each head is one-quarter of the original dimension.
[0037] Multi-scale pooling was performed on cross-modal interaction features to obtain a multi-scale feature set. Average pooling was used with kernel sizes of 2, 4, and 8, with the stride being the same as the kernel size and padding set to 0. Taking a 128-dimensional feature vector as an example, after pooling at three scales, 64-dimensional, 32-dimensional, and 16-dimensional features were obtained, forming the multi-scale feature set. Dynamic weight coefficients were calculated for each scale feature in the multi-scale feature set. The dynamic weight calculation employed a soft attention mechanism, calculating similarity scores between the learnable query vector and each scale feature. The similarity scores were normalized using a softmax function to obtain the weight coefficients. The query vector had a dimension of 32 and was initialized to a normal distribution with a mean of 0 and a standard deviation of 0.1. The multi-scale feature weights calculated for a steel plate scratch defect sample were: Scale 1 (64-dimensional) weight 0.45, Scale 2 (32-dimensional) weight 0.35, and Scale 3 (16-dimensional) weight 0.2, indicating that higher-dimensional features are more important for identifying this type of defect. The global fusion feature is obtained by adding the original cross-modal interaction features to a dynamically weighted multi-scale feature set. The weighting process is as follows: each scale feature is restored to its original dimension through linear upsampling, multiplied by the weight coefficients, summed, and then added to the original feature. Linear upsampling is implemented through a fully connected layer, with the number of parameters being the product of the feature dimensions.
[0038] Inter-layer skip connections are constructed to enhance information exchange between shallow and deep features. The output features of the previous layer are concatenated with the features of the current layer, and a fusion coefficient is calculated using a gating mechanism. Specifically, after concatenating the features of the previous and current layers, a two-layer fully connected network is used to generate gating coefficients. The number of neurons in the middle layer of the fully connected layer is half the dimension of the concatenated features, while the number of neurons in the output layer is the same as the feature dimension. The activation function is Sigmoid. Taking a 128-dimensional feature from the previous layer and a 128-dimensional feature from the current layer as an example, the concatenated feature dimension is 256, and the fully connected network structure is 256-128-128. The generated gating coefficient ranges from 0 to 1, representing the proportion of current layer features retained. In the detection of surface defects in aluminum alloy sheets, the average gating coefficient generated by the shallow network (layers 1 and 2) is 0.35, while the average gating coefficient of the deep network (layers 3 and 4) is 0.65, indicating that deeper features contribute more to defect detection. The updated features are obtained by dynamically fusing the features of the previous layer and the current layer based on the gating coefficient. The fusion method is as follows: the current layer feature is multiplied by the gating coefficient, the previous layer feature is multiplied by (1 - the gating coefficient), and the two are added together to obtain the updated features.
[0039] The mutual information between each layer's features and the target features is calculated to evaluate the contribution of each layer's features. The target features refer to the ideal feature representation corresponding to the defect labels during training, obtained through supervised learning. Mutual information is calculated using a kernel density estimation method, employing a Gaussian kernel function, with the bandwidth parameter set to the median of the sample variance. The range of mutual information is 0 to positive infinity; a larger value indicates a stronger correlation between the feature and the target. For example, the mutual information between the features of each layer and the target feature in a four-layer network structure are 0.42, 0.56, 0.75, and 0.61, respectively, indicating that the third layer's features have the strongest correlation with the target. The optimal number of network layers is determined based on the mutual information, and the network depth is dynamically adjusted during training to avoid overfitting and underfitting. The optimal number of layers is determined by stopping the addition of network layers when the increase in mutual information after adding a new layer is less than a preset threshold (set to 0.05). Adaptive weighting is applied to the features of each layer to obtain the final fused features. The weighting coefficients are proportional to the mutual information of each layer's features and are normalized to weight coefficients summing to 1 using the softmax function. Taking a four-layer network as an example, the generated weight coefficients are 0.18, 0.24, 0.32, and 0.26, respectively. The feature weighting method is as follows: after dimensionality unification processing, each layer's features are multiplied by their corresponding weight coefficients and then summed. Dimensionality unification is achieved through a projection matrix to ensure that all features are mapped to the same dimensional space (set to 160). The resulting final fused features serve as local correlation feature data, used for subsequent feature hypergraph structure construction and global correlation pattern fusion.
[0040] This invention achieves deep fusion and interaction of multimodal features through a multi-layered cascaded graph convolutional network structure, significantly improving the accuracy and robustness of surface defect detection in metallic materials. The hierarchical design of intramodal feature aggregation and intermodal feature interaction effectively captures complementary information and dependencies between different modalities. Multi-scale pooling and dynamic weighting mechanisms enhance the multi-scale expressive power of features, while the adaptive weighting strategy guided by inter-layer skip connections and mutual information optimizes the information flow path, avoiding information loss problems in deep networks.
[0041] Optionally, The step of constructing a feature hypergraph structure using preprocessed image data, acoustic feature data, electromagnetic feature data, and the locally associated feature data, wherein the feature hypergraph structure includes original modal feature nodes and locally associated feature nodes, and the hyperedges of the feature hypergraph structure represent the combination relationships between nodes includes: A node set is constructed using preprocessed image data, acoustic feature data, electromagnetic feature data, and local correlation feature data. An initial hyperedge set is generated based on the node set. The hyperedge weights are initialized using a multilayer perceptron. The initial node set is then weighted and aggregated to obtain the initial node representation. A mutual information matrix is constructed based on the initial node representation. A higher-order correlation tensor is constructed using the mutual information matrix. The node features are then enhanced based on the higher-order correlation tensor to obtain enhanced node features. Multi-scale pooling is performed on the enhanced node features to obtain a multi-layer feature pyramid. A feature mapping matrix between adjacent layers is constructed, and weighted fusion is performed based on the feature mapping matrix to obtain the fused feature representation. Cluster center nodes are determined by calculating node density and relative distance based on fused feature representation, a dynamic hyperedge structure is constructed, and the adaptive weights of the dynamic hyperedge structure are calculated. A temporal state transition function is constructed, and the node state and hyperedge structure are updated temporally based on the temporal state transition function. The updated nodes are mapped to knowledge graph entities to obtain knowledge features, and the knowledge features are fused with the node features to obtain an enhanced feature representation. A joint optimization objective is constructed based on reconstruction loss, structure preservation loss, and temporal consistency loss. The joint optimization objective is optimized to obtain the final dynamic adaptive hypergraph structure. The enhanced feature representation is input into the dynamic adaptive hypergraph structure for feature propagation to obtain the hypergraph structure feature representation. This hypergraph structure feature representation preserves the local structural relationships and combination patterns of multimodal data and is used for subsequent fusion with global associated features.
[0042] For example, when constructing the feature hypergraph structure, a node set is built using preprocessed image data, acoustic feature data, electromagnetic feature data, and local correlation feature data. Image data nodes are represented by 128-dimensional feature vectors, acoustic feature data nodes by 64-dimensional feature vectors, electromagnetic feature data nodes by 96-dimensional feature vectors, and local correlation feature data nodes by 160-dimensional feature vectors. Taking steel plate surface defect detection as an example, one detection sample generates 4 feature nodes. If the batch size is 32, a total of 128 nodes are generated. An initial hyperedge set is generated based on the node set, where each hyperedge represents a combination relationship of multiple nodes. Hyperedge generation uses a similarity clustering method, combining nodes with an Euclidean distance less than a threshold of 0.6 into a hyperedge. For a batch size of 32, approximately 48 hyperedges are generated, with each hyperedge containing an average of 3.5 nodes. The hyperedge weights are initialized using a three-layer multilayer perceptron, with each layer containing 128, 64, and 1 neurons respectively. The input is a concatenated vector of features from all nodes within the hyperedge, and the output is a scalar representing the hyperedge weights. The weight values range from 0 to 1, reflecting the importance of the hyperedge. In the aluminum plate surface scratch defect sample, the weight of the hyperedge containing image nodes and acoustic nodes is 0.85, and the weight of the hyperedge containing acoustic nodes and electromagnetic nodes is 0.78. The initial node representation is obtained by weighted aggregation of the initial node set. The aggregation method is as follows: the hyperedges connected to the nodes are weighted according to their weights, and then the features of other nodes within the hyperedges are aggregated. After aggregation, the feature dimension remains unchanged, but it includes local hypergraph structure information.
[0043] A mutual information matrix is constructed based on the initial node representation, describing the degree of information correlation between any two nodes. Mutual information is calculated using a kernel density estimation method, employing a radial basis function as the kernel function, with a bandwidth parameter set to 0.5. For a batch of 32 samples, the generated mutual information matrix has a size of 128×128, with element values ranging from 0 to 1. In a steel plate corrosion defect sample, the mutual information between image nodes and acoustic nodes is 0.72, between image nodes and electromagnetic nodes is 0.63, and between acoustic nodes and electromagnetic nodes is 0.85. A higher-order correlation tensor is constructed using the mutual information matrix to capture higher-order dependencies between multiple nodes. The higher-order correlation tensor is obtained through tensor product operations, with an order of 3 representing triplet relationships. The tensor size is 128×128×128. Considering computational efficiency, a low-rank approximation method is used to decompose the tensor into three low-dimensional factor matrices, each with a size of 128×32. Node features are enhanced using a higher-order correlation tensor. The enhancement method involves performing a shrunk operation between the tensor and the node features to obtain enhanced features that contain higher-order relationships. The enhanced features have the same dimension as the original features, but contain richer structural information.
[0044] Multi-scale pooling is performed on the enhanced node features to obtain a multi-layer feature pyramid. Graph pooling is used with a pooling ratio of 0.5, constructing a three-layer pyramid structure. Taking 128 nodes as an example, the three pyramid layers contain 128, 64, and 32 nodes respectively. Edge weights are used as pooling guidelines, prioritizing the retention of nodes with high weights. A feature mapping matrix between adjacent layers is constructed, describing the correspondence between upper-layer and lower-layer nodes. The mapping matrix is generated by calculating the cosine similarity between upper-layer and lower-layer nodes, and selecting the k lower-layer nodes with the highest similarity as the corresponding nodes of the upper-layer nodes, with k set to 3. The mapping matrix size is 64×128 between the first and second layers of the pyramid; and 32×64 between the second and third layers. Weighted fusion is performed based on the feature mapping matrix to obtain the fused feature representation. The fusion process involves weighted combination of upper-layer features and lower-layer features through the mapping matrix, with weight coefficients determined by learnable parameters. The initial weights are set to a uniform distribution and are automatically adjusted during training. The fused feature dimension is the same as the original node feature dimension, but it integrates multi-scale information.
[0045] Cluster center nodes are determined by calculating node density and relative distance based on fused feature representation. Node density is calculated by counting the number of nodes in a node's neighborhood, with the neighborhood radius set as the average Euclidean distance in the feature space. Relative distance is the minimum distance from a node to a node with higher density. Nodes with high density and large relative distance are selected as cluster centers, with the number of cluster centers adaptively determined, typically around 10% of the total number of nodes. In a copper plate surface defect detection task, 14 cluster center nodes were selected from 128 nodes. A dynamic hyperedge structure is constructed, forming a hyperedge with each cluster center node and its associated nodes. Associated nodes are determined using a distance threshold method, including nodes whose distance to the cluster center is less than a threshold of 0.7. Adaptive weights for the dynamic hyperedge structure are calculated, based on the consistency and diversity of node features within the hyperedge. Consistency is measured by the variance of node features, and diversity is measured by the entropy of node modality types. The weight calculation formula is: consistency score multiplied by diversity score, then normalized. These adaptive weights control the contribution of different hyperedges to information flow during hypergraph feature propagation, giving higher weights to hyperedges containing complementary information, while suppressing the influence of redundant or noisy hyperedges. In subsequent feature propagation stages, when nodes update features, they aggregate neighbor information according to the hyperedge weights, and the information transmitted by high-weight hyperedges will have a greater impact on the final feature representation.
[0046] A temporal state transition function is constructed to model the temporal evolution of node states and hyperedge structures. The state transition function is implemented using a gated recurrent unit (GRU) with a hidden layer dimension of 128. The input consists of the node features at the current time step and the state at the previous time step, and the output is the updated state. The initial state is initialized using a zero vector. The temporal update process is as follows: based on the current input and the previous state, an update gate and a reset gate are calculated using a gating mechanism, candidate states are generated, and finally, the new state is obtained by fusion. The updated nodes are mapped to entities in the knowledge graph to obtain knowledge features. The knowledge graph contains prior knowledge of defects in metallic materials, with 500 entities and 30 types of relationships. The mapping method is as follows: the similarity between node features and knowledge graph entity features is calculated, and the entity with the highest similarity is selected as the mapping target. The knowledge features and node features are fused to obtain an enhanced feature representation. The fusion method uses a gated fusion mechanism, adaptively adjusting the weights of the two features through learnable parameters. In a copper alloy surface corrosion detection sample, the node feature weight is 0.65 and the knowledge feature weight is 0.35, indicating that the model utilizes both data-driven features and prior knowledge.
[0047] A joint optimization objective is constructed based on reconstruction loss, structure preservation loss, and temporal consistency loss. Optimization of this objective yields the final dynamic adaptive hypergraph structure. The reconstruction loss measures the difference between the original features and the reconstructed hypergraph features, calculated using mean squared error. The structure preservation loss ensures that the hypergraph structure retains the topological relationships of the original data, implemented using a Laplacian feature map regularization term. The temporal consistency loss guarantees smooth changes in the hypergraph structure between adjacent time steps, measured by the difference in hyperedge weights between adjacent time steps. The weight ratio of the three losses is 3:2:1. The joint objective is optimized using gradient descent with a learning rate of 0.001 and 5000 iterations. The enhanced feature representation is input into the dynamic adaptive hypergraph structure for feature propagation, resulting in the hypergraph structure feature representation. The propagation process involves node features flowing along hyperedges; each node aggregates information from its connected hyperedges and updates its own features. After three rounds of propagation, the final feature representation is obtained. In the multi-defect detection task on aluminum plate surfaces, the features propagated through the hypergraph exhibit higher intra-class consistency and inter-class discriminability than the original features, providing a good foundation for subsequent fusion with globally related features.
[0048] The feature hypergraph structure constructed in this invention models the complex combination relationships between multimodal data through dynamic adaptive hyperedges, effectively integrating the advantages of local graph convolutional features and original modal features. The multi-layer feature pyramid structure and temporal state modeling mechanism enhance the multi-scale and temporal coherence of feature representations, while the knowledge graph-guided feature enhancement strategy incorporates prior knowledge into the data-driven model. The dynamic hyperedge adaptive generation and optimization method enables the hypergraph structure to flexibly adapt to different types of defect patterns, significantly improving the adaptability, robustness, and interpretability of the metal material surface defect detection system.
[0049] Optionally, The steps of calculating node density and relative distance based on fused feature representation to determine cluster center nodes, constructing dynamic hyperedge structures, and calculating the adaptive weights of the dynamic hyperedge structures include: Calculate the local density value and relative distance value between nodes, determine the node score based on the product of the local density value and the relative distance value between nodes, and determine the node center node if the node score is greater than the preset score threshold; based on the initial cluster center node, calculate the structural similarity and feature similarity of the node pairs, and obtain the comprehensive similarity by weighting the structural similarity and feature similarity. Based on the comprehensive similarity, nodes are grouped to obtain candidate hyperedges, the boundary nodes of the candidate hyperedges are identified, and the boundary nodes are added to the candidate hyperedges with the highest similarity to obtain optimized candidate hyperedges. An initial similarity threshold is determined based on the distribution of the comprehensive similarity. The initial similarity threshold is then adaptively adjusted in conjunction with the temporal stability of the hyperedge structure to obtain a final similarity threshold. The optimized candidate hyperedges are then selected based on the final similarity threshold. The degree of structural change is determined based on the node state changes and the hyperedge stability assessment results. Based on the degree of structural change, a corresponding update strategy is selected to dynamically adjust the hyperedge structure. For the adjusted hyperedge structure, the structural weight, feature weight, and temporal weight of the hyperedge are calculated. The final adaptive weight of the hyperedge is obtained by joint optimization based on local consistency, global structure preservation, and temporal smoothness.
[0050] Combination Figure 2The flowchart illustrating the construction of the dynamic hyperedge structure and the adaptive weight calculation is provided. For example, when calculating the local density values between nodes, a distance-based density estimation method is used to calculate the number of nodes in the neighborhood of each node. The neighborhood range is determined by a truncation distance parameter, which is set to the average distance between all nodes multiplied by 0.15. Taking a steel plate surface detection data as an example, the total number of nodes is 128, the feature dimension is 160, and the calculated truncation distance is 0.43. For each node, the number of neighboring nodes with a distance less than 0.43 in the feature space is used as its local density value. In practical applications, the average local density value of image modal nodes is 12.5, acoustic modal nodes is 9.3, electromagnetic modal nodes is 10.7, and locally associated feature nodes is 14.2. The relative distance between nodes is calculated, that is, the distance from each node to the nearest node with a density value greater than its own. For the node with the highest density, its relative distance is defined as the maximum distance between all nodes. In aluminum alloy surface defect detection, the relative distance values of a certain batch of data ranged from 0.25 to 1.32, with an average of 0.68. Node scores were calculated based on the product of the node's local density value and its relative distance value; this product reflects the node's suitability as a cluster center. A preset scoring threshold was set as the 75th percentile of all node scores, and nodes with scores above this threshold were identified as initial cluster center nodes. In a data batch of 128 nodes, the typical number of initial cluster center nodes ranged from 18 to 22.
[0051] Structural and feature similarities are calculated for all node pairs based on the initial cluster center node. The structural similarity calculation method involves constructing a structural descriptor for each node, containing the distance vector from the node to each cluster center, and calculating the cosine similarity between the structural descriptors of two nodes. In a sample of copper plate surface crack detection, the average structural similarity of nodes within the same defect region is 0.87, while the average structural similarity between nodes in different defect regions is 0.41. Feature similarity is calculated using the Euclidean distance between node feature vectors, with the distance inversely proportional to the similarity. The distance is converted to a similarity value using an exponential kernel function, with the kernel function bandwidth parameter set to 0.5. The average feature similarity between nodes of the same modality is 0.72, and the average feature similarity between nodes across modalities is 0.58. A weighted average similarity is obtained by weighting the structural and feature similarities. The weighting coefficients are determined through validation set optimization, with a structural similarity weight of 0.6 and a feature similarity weight of 0.4. This weighting method, while maintaining similarity in the feature space, emphasizes the positional relationship of nodes in the global structure, which is beneficial for discovering node groups with similar structural patterns.
[0052] Candidate hyperedges are obtained by grouping nodes based on comprehensive similarity. The grouping method uses density peak clustering, assigning nodes to the group corresponding to the cluster center with the highest comprehensive similarity, using the initial cluster center as the core. Initially, each group forms a candidate hyperedge. Boundary nodes of the candidate hyperedges are identified, defined as nodes with high similarity to multiple hyperedges. The criterion is that the difference between the average similarity of a node and its belonging hyperedge and the second most similar hyperedge is less than 0.15. In practical applications, an average of 18% of nodes are identified as boundary nodes. Boundary nodes are added to the candidate hyperedges with the highest similarity to obtain optimized candidate hyperedges. The redistribution of boundary nodes uses a soft allocation strategy, allowing a node to belong to multiple hyperedges simultaneously, with the membership degree proportional to the comprehensive similarity between the node and the hyperedge. In a sample of cracked aluminum plates, there are 32 boundary nodes, each assigned to an average of 2.3 hyperedges, with a membership degree threshold set to 0.25.
[0053] The initial similarity threshold is determined based on the distribution of comprehensive similarity. Histogram analysis is used to divide the comprehensive similarity value range [0, 1] into 20 equally wide intervals, and the frequency of each interval is counted. Significant valleys in the frequency distribution are identified as the initial threshold; if no significant valley is found, the median of the similarity is used. In defect detection of different types of metallic materials, the initial similarity threshold is typically between 0.45 and 0.65. The initial similarity threshold is adaptively adjusted based on the temporal stability of the hyperedge structure. Temporal stability is measured by the rate of change of the hyperedge structure in consecutive frames, calculated as the Hamming distance between the hyperedge structures of consecutive frames divided by the total number of nodes. An adjustment coefficient is set to 1 minus the temporal rate of change. When the rate of change is high, the similarity threshold is appropriately lowered to accommodate more edge nodes; when the rate of change is low, the threshold is increased to maintain structural stability. For example, in the detection of rapidly deformed regions on the aluminum plate surface, when the rate of change reaches 0.35, the threshold decreases from 0.60 to 0.39, allowing the 8 previously excluded edge nodes to be included in the hyperedge. However, in the detection of stable regions, the rate of change is only 0.05, and the threshold increases from 0.52 to 0.61, filtering out 12 weakly associated nodes and maintaining the stability of the hyperedge structure. Based on the final similarity threshold, optimized candidate hyperedges are selected, removing hyperedges with an average similarity between nodes below the threshold, and merging hyperedge pairs with a similarity above the threshold.
[0054] The degree of structural change is determined based on node state changes and hyperedge stability assessment results. Node state changes are calculated using the Euclidean distance of eigenvectors at continuous time steps, while hyperedge stability is measured by the proportion of changes in hyperedge member nodes. The degree of structural change is calculated by combining node state changes with hyperedge stability. When the degree of structural change is less than 0.2, an incremental update strategy is used; when the degree of change is between 0.2 and 0.6, a partial reconstruction strategy is used; and when the degree of change is greater than 0.6, a full reconstruction strategy is used. Incremental updates only adjust the allocation of boundary nodes; partial reconstruction preserves stable hyperedges and reconstructs changed hyperedges; and full reconstruction re-executes the entire hyperedge construction process. Three weights are calculated for the adjusted hyperedge structure. The structural weight reflects the importance of the hyperedge in the overall structure and is calculated as the average centrality of nodes within the hyperedge; this structural weight will be used to support global structure preservation constraints. The feature weight measures the consistency of features of nodes within the hyperedge and is calculated as the average cosine similarity of the feature vectors of nodes within the hyperedge; this feature weight will be used to support local consistency constraints. The temporal weight represents the stability of the hyperedge in the time dimension. It is calculated by the overlap ratio between the current hyperedge and the hyperedge at the previous time step. This temporal weight will be used to support temporal smoothness constraints.
[0055] The final adaptive weights of hyperedges are obtained through joint optimization based on local consistency, global structure preservation, and temporal smoothness. Specifically, the calculated feature weights ensure the similarity of node features within the hyperedge by supporting local consistency constraints; the calculated structure weights maintain the overall topology of the hypergraph by supporting global structure preservation constraints; and the calculated temporal weights limit drastic changes in hyperedge weights between adjacent time steps by supporting temporal smoothness constraints. The joint optimization objective function consists of three parts: consistency loss based on feature weights, structure loss based on structure weights, and smoothness loss based on temporal weights, with a weight ratio of 3:2:1. Gradient descent is used to optimize the joint objective, with a learning rate of 0.005 and 2000 iterations. The final adaptive weights of the hyperedges are distributed between 0.2 and 0.95, with a mean of 0.64 and a standard deviation of 0.17. These adaptive weights control the information flow during the hypergraph feature propagation process, with hyperedges having higher weights contributing more to feature updates. In a steel plate surface crack detection task, the average weight of hyperedges containing cross-modal nodes is 0.78, while the average weight of single-modal hyperedges is 0.52, indicating that the system places greater emphasis on hyperedges that fuse multimodal information. The weighted hypergraph structure can more effectively capture complementary information between different modes, thus improving feature representation capabilities.
[0056] This invention achieves high-order association modeling of multimodal data through dynamic construction of hypergraph structures and adaptive weight calculation. The hyperedge construction method based on density peak clustering and boundary node optimization improves the expressive power of the hypergraph structure for complex modal relationships. The adaptive weight calculation method with triple weight joint optimization balances multiple constraints of local consistency, global structure preservation, and temporal smoothness, enabling the hypergraph structure to efficiently and accurately express and transmit multimodal information.
[0057] Optionally, The steps for constructing a feature association tensor to capture global modal interaction patterns, and obtaining the global association features between modal features through tensor decomposition, include: Feature matrices are constructed from the preprocessed image data, acoustic feature data, and electromagnetic feature data, respectively. The image feature matrix, acoustic feature matrix, and electromagnetic feature matrix are then constructed into a third-order feature correlation tensor. The third-order feature correlation tensor is then standardized and sparsified to obtain a preprocessed tensor. The video-sound, video-electronic, and sound-electronic interaction matrices in the preprocessed tensor are calculated by tensor product operation. The norms of the three interaction matrices are calculated to obtain the local association strength. The local association strengths are aggregated to obtain the global association pattern. The preprocessed tensor is optimized based on the global association pattern. The optimized tensor is subjected to singular value decomposition. The optimal decomposition rank is determined based on the energy retention rate and a preset energy threshold. The optimal decomposition rank is used to decompose the tensor to obtain a group of factor vectors. The importance weight of the pattern is calculated based on the norm of the factor vectors. The group of factor vectors is then weighted and combined to obtain the global association feature. The norm of the global association feature is calculated to obtain the regularization constraint, and the difference norm between the global association feature and the local association strength is calculated to obtain the consistency constraint. The regularization constraint and the consistency constraint are weighted and combined to construct the feature optimization objective. The feature optimization objective is optimized to obtain the optimized global association feature. The feature difference norm is calculated between the current global correlation feature and the global correlation feature at the previous time step to obtain the feature difference degree. When the feature difference degree is greater than the adaptive update threshold, the current global correlation feature and the global correlation feature at the previous time step are weighted and fused to update the global correlation feature.
[0058] For example, feature matrices are constructed for the preprocessed image data, acoustic feature data, and electromagnetic feature data, respectively. Rows in the feature matrices represent samples, and columns represent feature dimensions. Taking the detection of surface defects in metallic materials as an example, the image feature matrix has a dimension of batch size × 128, the acoustic feature matrix has a dimension of batch size × 64, and the electromagnetic feature matrix has a dimension of batch size × 96. When the batch size is 32, the dimensions of these three matrices are 32 × 128, 32 × 64, and 32 × 96, respectively. These three feature matrices are constructed into a third-order feature correlation tensor using direct outer product operations, resulting in a tensor dimension of 128 × 64 × 96. Each element represents the correlation strength of the corresponding three modal features. The third-order feature correlation tensor is standardized by subtracting the mean of each element and dividing by the standard deviation, resulting in a tensor element distribution with a mean of 0 and a standard deviation of 1. The standardized tensor is then sparsified by retaining elements with an absolute value greater than 0.1 and setting the rest to 0.
[0059] Modal interaction matrices in the preprocessed tensor were calculated using tensor product operations, including three interaction matrices: visual-acoustic, visual-electromagnetic, and acoustic-electromagnetic. The tensor product operation method involved compressing and summing the tensor along the third dimension to obtain the visual-acoustic interaction matrix (128×64). Similarly, compression along the second dimension yielded the visual-electromagnetic interaction matrix (128×96), and compression along the first dimension yielded the acoustic-electromagnetic interaction matrix (64×96). The norms of the three interaction matrices were calculated to determine the local correlation strength. The Frobenius norm, the square root of the sum of the squares of all elements in the matrix, was used for norm calculation. In the aluminum alloy surface corrosion defect detection samples, the norm of the visual-acoustic interaction matrix was 42.3, the norm of the visual-electromagnetic interaction matrix was 38.7, and the norm of the acoustic-electromagnetic interaction matrix was 46.2, indicating the strongest interaction between acoustic and electromagnetic features. The global association pattern is obtained by aggregating the local association strengths using a weighted average method with a weight ratio of visual-acoustic:visual-electrical:acoustic-electrical = 1:1:1.2, slightly emphasizing acoustic-electrical interaction. The aggregated global association pattern is a scalar value. Among different types of defect samples, the average global association pattern value is 82.6 for crack defects, 75.3 for corrosion defects, and 69.8 for pit defects. The preprocessed tensor is optimized based on the global association pattern. The optimization method is to use the ratio of tensor elements to global association pattern values as weights to weight the tensor.
[0060] Singular value decomposition (SVD) was performed on the optimized tensor using a high-order SVD algorithm. The optimal decomposition rank was determined based on the energy retention rate and a preset energy threshold. The energy retention rate was defined as the proportion of the sum of squares of the first k singular values to the total sum of squares of singular values. The preset energy threshold was set to 0.85, meaning 85% of the energy was retained. Tensor decomposition using the optimal decomposition rank yielded factor vector groups, containing three groups of factor vectors corresponding to image, acoustic, and electromagnetic modes, respectively, with each group containing 16 factor vectors. The image mode factor vector had a dimension of 128, the acoustic mode factor vector had a dimension of 64, and the electromagnetic mode factor vector had a dimension of 96. The mode importance weights were calculated based on the factor vector norm, using the L2 norm of the vector. In a sample of copper plate surface defect detection, the importance weight of the first mode was 0.21, the second was 0.18, and the third was 0.15, decreasing sequentially. The global association feature is obtained by weighting and combining the factor vectors. The combination method is to take the outer product of the factor vectors at corresponding positions of the three modalities, multiply them by the importance weights, and then sum them. The resulting global association feature has a dimension of 128×64×96. In practical applications, to reduce computational complexity, it can be reshaped into a one-dimensional vector with a dimension of 128×64×96.
[0061] The regularization constraint is obtained by calculating the norm of the global association features, using the L2 norm. The consistency constraint is obtained by calculating the difference norm between the global association features and the local association strengths, calculated as the Euclidean distance between the interaction matrix norm obtained from the tensor product operation of the global association features and the previously calculated local association strengths. The regularization constraint and the consistency constraint are weighted and combined to construct the feature optimization objective, with a weight ratio of regularization constraint:consistency constraint = 0.3:0.7, emphasizing the consistency between features and local association patterns. Gradient descent is used for optimization, with a learning rate of 0.01 and 1000 iterations. Optimizing the feature optimization objective yields the optimized global association features.
[0062] The feature dissimilarity is calculated by taking the norm of the difference between the current global associated features and the global associated features from the previous time step. The norm of the difference is the cosine distance, which is 1 minus the cosine similarity between the two feature vectors. The adaptive update threshold is set by adding 0.5 times the standard deviation of the historical feature dissimilarity, with an initial threshold of 0.2, which is continuously updated during the detection process. When the feature dissimilarity exceeds the adaptive update threshold, the current global associated features and the global associated features from the previous time step are weighted and fused to update the global associated features. The fusion weight is positively correlated with the feature dissimilarity, and the calculation formula is: the current feature weight equals the base weight 0.6 plus 0.4 multiplied by the ratio of the dissimilarity to the threshold. In practical applications, when the feature difference is 0.25 and the threshold is 0.2, the current feature weight is 0.6 + 0.4 × (0.25 / 0.2) = 1.1. Since the weight exceeds 1, it is taken as 1, so the weight is 1, meaning the current feature is used completely. When the feature difference is 0.22 and the threshold is 0.2, the current feature weight is 0.6 + 0.4 × (0.22 / 0.2) = 1.04, meaning the weight is 1. When the feature difference is 0.15 and the threshold is 0.2, the current feature weight is 0.6 + 0.4 × (0.15 / 0.2) = 0.9, and the feature weight at the previous time step was 0.1.
[0063] In the implementation of tensor decomposition, alternating least squares optimization is employed. The core idea of alternating least squares is to fix other factors while optimizing the current factor. Taking third-order tensor decomposition as an example, the alternating optimization process consists of three steps: fixing acoustic and electromagnetic factors and optimizing image factors; fixing image and electromagnetic factors and optimizing acoustic factors; and fixing image and acoustic factors and optimizing electromagnetic factors. Each optimization step uses least squares to solve a system of linear equations. To improve decomposition efficiency, random initialization is used in the actual implementation to accelerate convergence. The initialization method involves random sampling from a normal distribution with a mean of 0 and a standard deviation of 0.01. When calculating global correlation features, to avoid the curse of dimensionality, Khatri-Rao product operations are used to directly calculate the low-dimensional representation instead of explicitly constructing a high-dimensional tensor. The low-dimensional representation of the global correlation features has a dimension of 16×(128+64+96)=4624, which is approximately 170 times lower than the original dimension of 128×64×96=786432, significantly improving computational efficiency.
[0064] To prevent feature instability caused by frequent updates during the temporal update process of globally correlated features, a momentum mechanism is introduced. The momentum parameter is set to 0.9, and the update formula is: the new feature equals the momentum parameter multiplied by the old feature plus (1 - momentum parameter) multiplied by the currently calculated feature. In addition to using cosine distance, the feature dissimilarity calculation also incorporates Mahalanobis distance in the feature space, better measuring the changes in feature statistical distribution. Mahalanobis distance is calculated using the inverse of the feature covariance matrix, which is estimated from historical feature samples. The two distances are weighted and combined in a 6:4 ratio to form the final feature dissimilarity. In the detection of surface cracks on aluminum plates, the false alarm rate of pure cosine distance is 8%, which decreases to 3.5% after combining with Mahalanobis distance, indicating that the combined measure is more robust.
[0065] This invention effectively captures global interaction patterns among multimodal data by constructing a third-order feature association tensor and tensor decomposition techniques. An optimal decomposition rank determination strategy guided by energy preservation rate and a factor vector importance weighting mechanism ensure that the tensor decomposition results retain key information. The joint optimization objective of regularization and consistency constraints, along with an adaptive feature update mechanism, improves the robustness and temporal coherence of globally associated features.
[0066] Finally, the hypergraph structural feature representation and the global association feature are fused. In the feature fusion stage, the hypergraph structural feature representation and the global association feature are used as inputs. The hypergraph structural feature representation has a dimension of batch size × 256, representing the local structural association information of the multimodal data; the global association feature has a dimension of batch size × 512, containing the interaction patterns between modalities. For example, when the batch size is 32, the hypergraph structural feature representation has a dimension of 32 × 256, and the global association feature has a dimension of 32 × 512. The correlation matrix between these two feature representations is calculated to obtain the cross-attention weights. The calculation method is to perform matrix multiplication between the hypergraph features and the global features, and then normalize using the Softmax function. Specifically, the transpose of the hypergraph features and the global features is multiplied to obtain the correlation matrix, which has a dimension of 256 × 512. Then, the Softmax function is applied to each row to make the sum equal to 1. In the steel plate surface crack detection task, the average maximum value in the correlation matrix is 0.087, and the average minimum value is 0.0005, indicating that there are significant differences in the correlation between features.
[0067] Cross-attention weights are used to mutually enhance hypergraph structural feature representations and global correlation features, with the enhancement process involving bidirectional information flow. The hypergraph structural features enhance the global correlation features by multiplying the correlation matrix with the hypergraph features to obtain an enhanced matrix of dimension 512×256, which is then added to the original global features. Similarly, the global correlation features enhance the hypergraph structural features by multiplying the transpose of the correlation matrix with the global features to obtain an enhanced matrix of dimension 256×512, which is then added to the original hypergraph features.
[0068] The enhanced features are fused with the original features via residual connections, which are performed by direct addition. The fusion formula for hypergraph structural features is: enhanced hypergraph features plus original hypergraph features, multiplied by 0.5; the fusion formula for global association features is similar. This residual connection mechanism effectively mitigates the information loss problem that may occur during feature enhancement. A weighted combination of feature consistency loss and complementarity loss is constructed as the fusion optimization objective. The consistency loss is calculated as the cosine distance between the fused hypergraph features and the fused global features, and the complementarity loss is calculated as the negative value of the orthogonality of the two features. In the case of steel pipe surface scratch detection, the consistency loss weight is set to 0.65, and the complementarity loss weight is set to 0.35, emphasizing the consistency of feature representation. The final fused features are obtained by optimizing the fusion optimization objective using gradient descent with a learning rate of 0.005 and 500 iterations. The final fused feature dimension is batch size × 384, which is 32 × 384 in practical applications.
[0069] The final fused features are input into a classifier, which employs a fully connected layer structure with three hidden layers containing 256, 128, and 64 nodes respectively, and using ReLU activation. The classifier compares the features to a pre-defined defect type feature library, which includes standard feature representations of typical metal surface defects such as cracks, corrosion, pits, and scratches, with 50 sample features for each defect type. The defect type and confidence level are obtained based on feature similarity. Cosine similarity is used for similarity calculation, and the defect type corresponding to the highest similarity is taken as the detection result, with the similarity value serving as the confidence level. For example, a similarity threshold of 0.75 is set; detection results below this threshold are marked as "unknown defects." The output defect detection results include defect type, confidence level, location information, and area estimation. The detection results include the following fields: defect ID, type, confidence level, coordinates, size, and timestamp.
[0070] Secondly, a multi-sensor-based surface defect detection system for metallic materials is provided, including: The first unit is used to acquire image data, acoustic feature data and electromagnetic feature data of the surface of the metal material through optical sensors, ultrasonic sensors and eddy current sensors respectively, and to perform preprocessing on each data. The second unit is used to input the preprocessed image data, acoustic feature data and electromagnetic feature data into the graph convolution feature fusion network. The three modal features are constructed into a local correlation graph through a multi-layer graph convolution structure. The nodes in the local correlation graph represent the feature vectors of each modality, and the edges represent the local dependencies between adjacent features. The local correlation feature data is obtained through cascaded graph convolution operations. The third unit is used to construct a feature hypergraph structure using preprocessed image data, acoustic feature data, electromagnetic feature data, and the local correlation feature data. The feature hypergraph structure includes original modal feature nodes and local correlation feature nodes, and the hyperedges of the feature hypergraph structure represent the combination relationship between nodes. A feature correlation tensor is constructed to capture the global modal interaction pattern, and the global correlation features between modal features are obtained through tensor decomposition. The feature hypergraph structure and the global correlation features are adaptively weighted and fused based on an attention mechanism to output the defect detection result.
[0071] Thirdly, a computer-readable storage medium is provided, having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
Claims
1. A method for detecting surface defects in metallic materials based on multiple sensors, characterized in that, include: Image data, acoustic feature data, and electromagnetic feature data of the metal material surface are acquired by optical sensors, ultrasonic sensors, and eddy current sensors, respectively, and preprocessed accordingly. Preprocessed image data, acoustic feature data, and electromagnetic feature data are input into a graph convolutional feature fusion network. The three modal features are constructed into a local correlation graph through a multi-layer graph convolutional structure. In the local correlation graph, nodes represent the feature vectors of each modality, and edges represent the local dependencies between adjacent features. Local correlation feature data is obtained through cascaded graph convolution operations. A feature hypergraph structure is constructed using preprocessed image data, acoustic feature data, electromagnetic feature data, and the local correlation feature data. The feature hypergraph structure includes original modal feature nodes and local correlation feature nodes. The hyperedges of the feature hypergraph structure represent the combination relationship between nodes. A feature association tensor is constructed to capture global modal interaction patterns, and global association features between modal features are obtained through tensor decomposition. An attention mechanism is used to adaptively weight and fuse the feature hypergraph structure and the global associated features to output the defect detection result.
2. The method according to claim 1, characterized in that, The steps for acquiring image data, acoustic feature data, and electromagnetic feature data of a metallic material surface using optical sensors, ultrasonic sensors, and eddy current sensors, respectively, include: The grayscale variance and signal-to-noise ratio of the image data, the spectral kurtosis and energy dispersion of the acoustic feature data, and the impedance change rate and phase fluctuation of the electromagnetic feature data are calculated and combined to obtain feature statistics. The feature statistics are converted into membership values using a Gaussian membership function, and the uncertainty score is calculated based on the membership values. The secondary sampling region is determined based on the spatial distribution characteristics of the uncertainty score. When the uncertainty score exceeds the preset threshold of the corresponding sensor, parameter adjustment is triggered, wherein: the imaging resolution of the optical sensor adopts an exponential adjustment function; the frequency bandwidth and pulse repetition frequency of the ultrasonic sensor adopt a linear adjustment function; the excitation frequency and scanning density of the eddy current sensor adopt a logarithmic adjustment function; the independent variable of the adjustment function is the uncertainty score, and the dependent variable is the ratio of the adjusted parameter to the initial parameter; the adjusted parameters are used to collect data in the secondary sampling area.
3. The method according to claim 1, characterized in that, The steps of inputting preprocessed image data, acoustic feature data, and electromagnetic feature data into a graph convolutional feature fusion network to obtain local correlation feature data include: Feature extraction is performed on the preprocessed image data, acoustic feature data, and electromagnetic feature data to obtain image feature vectors, acoustic feature vectors, and electromagnetic feature vectors, respectively. The Gaussian similarity matrix between the image feature vectors, acoustic feature vectors, and electromagnetic feature vectors is calculated. An adjacency matrix is constructed based on the Gaussian similarity matrix, and the adjacency matrix is normalized to obtain a normalized adjacency matrix. A multi-layered cascaded graph convolutional network structure is constructed, comprising an intra-modal feature aggregation layer, a cross-modal feature interaction layer, and a global feature fusion layer. Graph convolution is performed by multiplying the normalized adjacency matrix with the input feature set and the learnable weight matrix. The intra-modal feature aggregation layer uses the normalized adjacency matrix to perform graph convolution on a single modality feature. The cross-modal feature interaction layer uses the normalized adjacency matrix to interact and fuse features from different modalities. The global feature fusion layer uses the normalized adjacency matrix to integrate multi-modal features into a unified representation. A channel attention mechanism is introduced into the graph convolutional network structure. The importance score of the feature channel is calculated through a learnable weight matrix, and the feature is weighted and enhanced according to the importance score of the feature channel. The output features of the graph convolutional network structure are subjected to feature fusion processing to obtain local correlation feature data.
4. The method according to claim 3, characterized in that, The steps for constructing a multi-layered cascaded graph convolutional network structure include: Image feature vectors, acoustic feature vectors, and electromagnetic feature vectors are input into the intra-modal feature aggregation layer. The intra-modal weight vector is calculated by a multilayer perceptron. The image feature vectors, acoustic feature vectors, and electromagnetic feature vectors are weighted and aggregated based on the intra-modal weight vectors and the normalized adjacency matrix to obtain the intra-modal features. Intermodal attention maps are obtained by calculating normalized inner products between different modal features. Based on the intermodal attention maps, the intramodal features are mapped to value features. The value features are then weighted and combined to obtain cross-modal interaction features. The cross-modal interaction features are pooled at multiple scales to obtain a multi-scale feature set. The dynamic weights of each scale feature in the multi-scale feature set are calculated. The cross-modal interaction features are added to the multi-scale feature set weighted by the dynamic weights to obtain a global fusion feature. Construct inter-layer skip connections, concatenate the output features of the previous layer with the features of the current layer to calculate the gating coefficient, and dynamically fuse the output features of the previous layer and the features of the current layer based on the gating coefficient to obtain updated features; Calculate the mutual information between the features of each layer and the target features, determine the optimal number of network layers based on the mutual information, and adaptively weight the features of each layer to obtain the final fused features.
5. The method according to claim 1, characterized in that, The step of constructing a feature hypergraph structure using preprocessed image data, acoustic feature data, electromagnetic feature data, and the locally associated feature data, wherein the feature hypergraph structure includes original modal feature nodes and locally associated feature nodes, and the hyperedges of the feature hypergraph structure represent the combination relationships between nodes includes: A node set is constructed using preprocessed image data, acoustic feature data, electromagnetic feature data, and local correlation feature data. An initial hyperedge set is generated based on the node set. The hyperedge weights are initialized using a multilayer perceptron. The initial node set is then weighted and aggregated to obtain the initial node representation. A mutual information matrix is constructed based on the initial node representation. A higher-order correlation tensor is constructed using the mutual information matrix. The node features are then enhanced based on the higher-order correlation tensor to obtain enhanced node features. Multi-scale pooling is performed on the enhanced node features to obtain a multi-layer feature pyramid. A feature mapping matrix between adjacent layers is constructed, and weighted fusion is performed based on the feature mapping matrix to obtain the fused feature representation. Cluster center nodes are determined by calculating node density and relative distance based on fused feature representation, a dynamic hyperedge structure is constructed, and the adaptive weights of the dynamic hyperedge structure are calculated. A temporal state transition function is constructed, and the node state and hyperedge structure are updated temporally based on the temporal state transition function. The updated nodes are mapped to knowledge graph entities to obtain knowledge features, and the knowledge features are fused with the node features to obtain an enhanced feature representation. A joint optimization objective is constructed based on reconstruction loss, structure preservation loss, and temporal consistency loss. The joint optimization objective is then optimized to obtain the final dynamic adaptive hypergraph structure. The enhanced feature representation is then input into the dynamic adaptive hypergraph structure for feature propagation.
6. The method according to claim 5, characterized in that, The steps of calculating node density and relative distance based on fused feature representation to determine cluster center nodes, constructing dynamic hyperedge structures, and calculating the adaptive weights of the dynamic hyperedge structures include: Calculate the local density value and relative distance value between nodes, determine the node score based on the product of the local density value and the relative distance value between nodes, and determine the node center node if the node score is greater than the preset score threshold; based on the initial cluster center node, calculate the structural similarity and feature similarity of the node pairs, and obtain the comprehensive similarity by weighting the structural similarity and feature similarity. Based on the comprehensive similarity, nodes are grouped to obtain candidate hyperedges, the boundary nodes of the candidate hyperedges are identified, and the boundary nodes are added to the candidate hyperedges with the highest similarity to obtain optimized candidate hyperedges. An initial similarity threshold is determined based on the distribution of the comprehensive similarity. The initial similarity threshold is then adaptively adjusted in conjunction with the temporal stability of the hyperedge structure to obtain a final similarity threshold. The optimized candidate hyperedges are then selected based on the final similarity threshold. The degree of structural change is determined based on the node state changes and the hyperedge stability assessment results. Based on the degree of structural change, a corresponding update strategy is selected to dynamically adjust the hyperedge structure. For the adjusted hyperedge structure, the structural weight, feature weight, and temporal weight of the hyperedge are calculated. The final adaptive weight of the hyperedge is obtained by joint optimization based on local consistency, global structure preservation, and temporal smoothness.
7. The method according to claim 1, characterized in that, The steps for constructing a feature association tensor to capture global modal interaction patterns, and obtaining the global association features between modal features through tensor decomposition, include: The image feature matrix, acoustic feature matrix, and electromagnetic feature matrix are constructed into a third-order feature correlation tensor. The third-order feature correlation tensor is then standardized and sparsified to obtain a preprocessed tensor. The video-sound, video-electronic, and sound-electronic interaction matrices in the preprocessed tensor are calculated by tensor product operation. The norms of the three interaction matrices are calculated to obtain the local association strength. The local association strengths are aggregated to obtain the global association pattern. The preprocessed tensor is optimized based on the global association pattern. The optimized tensor is subjected to singular value decomposition. The optimal decomposition rank is determined based on the energy retention rate and a preset energy threshold. The optimal decomposition rank is used to decompose the tensor to obtain a group of factor vectors. The importance weight of the pattern is calculated based on the norm of the factor vectors. The group of factor vectors is then weighted and combined to obtain the global association feature. The norm of the global association feature is calculated to obtain the regularization constraint, and the difference norm between the global association feature and the local association strength is calculated to obtain the consistency constraint. The regularization constraint and the consistency constraint are weighted and combined to construct the feature optimization objective, and the optimized global association feature is obtained. The feature difference norm is calculated between the current global correlation feature and the global correlation feature at the previous time step to obtain the feature difference degree. When the feature difference degree is greater than the adaptive update threshold, the current global correlation feature and the global correlation feature at the previous time step are weighted and fused to update the global correlation feature.
8. A multi-sensor-based surface defect detection system for metallic materials, used to implement the method of any one of claims 1-7, characterized in that, include: The first unit is used to acquire image data, acoustic feature data and electromagnetic feature data of the surface of the metal material through optical sensors, ultrasonic sensors and eddy current sensors respectively, and to perform preprocessing on each data. The second unit is used to input the preprocessed image data, acoustic feature data and electromagnetic feature data into the graph convolution feature fusion network. The three modal features are constructed into a local correlation graph through a multi-layer graph convolution structure. The nodes in the local correlation graph represent the feature vectors of each modality, and the edges represent the local dependencies between adjacent features. The local correlation feature data is obtained through cascaded graph convolution operations. The third unit is used to construct a feature hypergraph structure using preprocessed image data, acoustic feature data, electromagnetic feature data, and the local associated feature data. The feature hypergraph structure includes original modal feature nodes and local associated feature nodes. The hyperedges of the feature hypergraph structure represent the combination relationship between nodes. A feature association tensor is constructed to capture global modal interaction patterns, and global association features between modal features are obtained through tensor decomposition. An attention mechanism is used to adaptively weight and fuse the feature hypergraph structure and the global associated features to output the defect detection result.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.
Citation Information
Cited By
Multi-modal dynamic fusion method based on graph attention network low-rank decomposition
CN121527584A
A metal bar quality detection method based on ultrasonic flaw detection
CN122385762A