Nondestructive testing and identification method for high-strength steel welding defects based on multi-modal data fusion
By using a multimodal data fusion method that combines two-dimensional images and three-dimensional point cloud data, the problem of the inability of two-dimensional images to quantify the three-dimensional size of welding defects was solved, enabling accurate identification and quantitative analysis of welding defects in high-strength steel, and improving the accuracy and reliability of detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI CONSTRUCTION GROUP CO LTD
- Filing Date
- 2025-12-03
- Publication Date
- 2026-04-21
AI Technical Summary
Existing two-dimensional image recognition methods cannot accurately quantify the three-dimensional dimensions of welding defects, and are prone to false detection or missed detection due to factors such as light, oil stains, and spatter.
A multimodal data fusion method is adopted, combining two-dimensional images and three-dimensional point cloud data. Through spatial alignment, feature extraction and fusion, a deep learning network is used to identify and classify defects, and obtain the three-dimensional size and category of defects.
It enables accurate identification and precise quantitative analysis of welding defects, effectively suppressing false detections and missed detections, and providing an objective assessment of the welding quality of high-strength steel.
Smart Images

Figure CN121236080B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of welding technology, and in particular to a non-destructive testing and identification method for welding defects in high-strength steel based on multimodal data fusion. Background Technology
[0002] High-strength steel, due to its excellent strength-to-weight ratio and good mechanical properties, has been widely used in key industrial fields such as shipbuilding, aerospace, bridge construction, and pressure vessels. Welding is one of the most important connection processes in high-strength steel structures, and the quality of welding directly determines the safety and service life of the structure. However, the welding process is easily affected by process parameters, environmental factors, and the characteristics of the material itself, and is prone to various types of defects such as porosity, cracks, lack of fusion, undercut, and weld beads. These defects can significantly reduce the fatigue strength and service safety of the structure, so accurate and reliable non-destructive testing and quality assessment of welded joints are crucial. To perform non-destructive testing on welding defects in high-strength steel, an automated welding defect identification method based on two-dimensional images is used in this field. This method acquires two-dimensional images of the weld surface using an industrial camera and uses image processing algorithms or deep learning models to identify and classify defects. Its disadvantages are that, firstly, two-dimensional images only have two dimensions, length and width, and lack the third dimension of geometric information, namely depth, which leads to the problem of not being able to accurately quantify the three-dimensional dimensions of welding defects (such as depth and volume), such as the depth of porosity, the amount of undercut, or the protrusion height of weld beads. Secondly, for defects with subtle surface texture changes but three-dimensional morphological features, or texture interference caused by light, oil, or splashes, relying solely on two-dimensional image information can easily lead to false detections or missed detections. Summary of the Invention
[0003] The purpose of this invention is to provide a non-destructive testing and identification method for high-strength steel welding defects based on multimodal data fusion, so as to solve the problems that two-dimensional image recognition of welding defects cannot accurately quantify the three-dimensional size information of welding defects and that two-dimensional image recognition of welding defects is prone to false detection or missed detection.
[0004] To address the aforementioned technical problems, the present invention provides a method for non-destructive testing and identification of welding defects in high-strength steel based on multimodal data fusion, comprising:
[0005] Welding data acquisition: Acquire two-dimensional images and three-dimensional point cloud data of the weld joint of high-strength steel to form data pairs;
[0006] Data spatial alignment: Establish the mapping relationship between the pixels of the two-dimensional image and the three-dimensional point cloud data, form a depth map corresponding to the two-dimensional image, and generate spatially aligned image-depth map data pairs;
[0007] Feature extraction and fusion: Based on the 2D image-depth map data, feature maps and 3D feature volumes of the 2D image are extracted to obtain 2D visual features and 3D feature volumes with the same number of feature channels. The 3D feature volume is used to guide the attention of the 2D visual features to generate a spatial attention weight map. The spatial attention weight map is used to weight the 2D visual features with geometric information and then fused with the 3D feature volume to obtain fused features with spatial geometric information and 2D visual information.
[0008] Defect identification and classification: The decoder restores the spatial dimensions of the fused features and associates the feature information with the pixel position. The identified defective pixels are quantized to calculate the three-dimensional dimensions of the defects. The feature information and three-dimensional dimensions associated with the connected components of each defective pixel are used as input. A pre-trained fully connected neural network classifier is used to automatically classify the defect categories.
[0009] Furthermore, the non-destructive testing and identification method for high-strength steel welding defects based on multimodal data fusion provided by the present invention includes the following welding data acquisition method:
[0010] Arrange reference points: Set no fewer than four visual reference points symmetrically on the base material on both sides of the weld after high-strength steel welding;
[0011] Deploy acquisition equipment: Set up a defect detection path, distribute multiple detection stations along the defect detection path, and deploy acquisition equipment at the detection stations, wherein the acquisition equipment integrates an industrial camera and a 3D scanner.
[0012] Data acquisition and matching: Two-dimensional images of the weld are acquired using an industrial camera, and three-dimensional point cloud data of the weld are acquired using a 3D scanner. The acquisition equipment then associates and stores the two-dimensional images and three-dimensional point cloud data to form data pairs.
[0013] Furthermore, the non-destructive testing and identification method for high-strength steel welding defects based on multimodal data fusion provided by the present invention includes the following data spatial alignment method:
[0014] 2D Image Reference Point Detection: Reference points in a 2D image are identified using image recognition technology. Subpixel-level edge extraction is performed on the identified reference points, and the geometric center pixel coordinates are calculated through ellipse fitting. ,in and These are pixel coordinate values. ,in i As a reference point, N A positive integer, which represents the coordinates of a reference point in the two-dimensional image as the base point;
[0015] 3D point cloud data reference point detection: The acquired 3D point cloud data is subjected to pass-through filtering to retain the 3D point cloud data near the weld seam. The data near the weld seam is then filtered to remove outlier noise points. For the filtered 3D point cloud data, the normal vector and curvature of each point are calculated. Threshold segmentation is used to initially extract the point cloud data corresponding to the reference points. A distance-based clustering algorithm is then used to group the initially extracted 3D point cloud data, with each cluster representing a reference point. The 3D coordinates of the center point of each cluster are calculated. This yields the three-dimensional coordinates of the reference point's position, where , and These are spatial coordinate values. ,in i The number of reference points, N It is a positive integer.
[0016] Solve for the coordinate transformation matrix:
[0017] Using the pixel coordinates of the two-dimensional image as a reference, the three-dimensional coordinates are projected onto the two-dimensional pixel coordinates according to the coordinate transformation matrix:
[0018] (1);
[0019] In equation (1), It is a scaling factor. These are pixel coordinates. These are the three-dimensional coordinates of the point cloud data; It is the camera intrinsic parameter matrix. yes The rotation matrix, yes The translation vector, It is the camera extrinsic parameter matrix;
[0020] According to formula (1), the matching point pairs and the known camera intrinsic parameter matrix As input, solve for the camera extrinsic matrix. .
[0021] Data spatial alignment:
[0022] Based on the projection matrix, the 3D point cloud data is projected onto the 2D image pixel coordinate system:
[0023] (4);
[0024] In equation (4), These are the theoretical pixel coordinates corresponding to the point cloud data. Represents the equivalence relation in homogeneous coordinates. It is a projection matrix. The three-dimensional coordinates of the point cloud data;
[0025] All projected points are determined based on their theoretical pixel coordinates and the original... Value resampling generates a depth map with the same resolution as the two-dimensional image, forming an aligned image-depth map data pair.
[0026] Furthermore, the non-destructive testing and identification method for high-strength steel welding defects based on multimodal data fusion provided by the present invention includes the following feature extraction and fusion methods:
[0027] Two-dimensional image feature extraction:
[0028] Using a 2D image from an image-depth map data pair as input, a pre-trained deep convolutional neural network (CNN) is used as the backbone to extract feature information from the 2D image, including edges and textures. After the input 2D image propagates forward through the backbone network, a high-dimensional feature map is obtained in the last convolutional layer. , It is The three-dimensional tensor, in which and These represent the height and width of the feature map, respectively. It is the number of feature channels. The visual appearance information of the weld area is encoded;
[0029] 3D point cloud feature extraction:
[0030] Voxelization of 3D point cloud data is performed to form a feature voxel mesh;
[0031] The feature voxel mesh is input into a 3D convolutional neural network (3DCNN) to extract the 3D geometric feature information of points in the point cloud data. This 3D geometric feature information includes convexity, curvature, and depth variations. After forward propagation processing by the 3DCNN, the 3D feature volume is obtained. , It is The four-dimensional tensor, in which , and These are the depth, height, and width of the feature volume, respectively. It is the number of channels for the three-dimensional feature;
[0032] Multimodal feature fusion:
[0033] Three-dimensional feature volume Spatial dimensions and feature diagrams Align to obtain dimensions of Two-dimensional visual features of three-dimensional tensors and three-dimensional feature volume Using three-dimensional feature volume To guide two-dimensional visual features Attention generation spatial attention weight map That is, a spatial attention weight map is generated through a lightweight subnetwork. :
[0034] (6);
[0035] In equation (6), A represents the spatial attention weight map. This refers to the Sigmoid function, used to compress weights to... interval; and They are respectively The weights and biases of convolution. This represents the convolution operation. It is a three-dimensional feature volume;
[0036] Using the spatial attention weight map Two-dimensional visual features Perform geometric information weighting and combine it with three-dimensional feature volume. Feature fusion is performed to obtain fused features containing both spatial geometric information and two-dimensional visual information:
[0037] (7);
[0038] In equation (7), As a feature of fusion, and They are respectively The weights and biases of convolution. This represents the convolution operation. For the weighted features, Indicates will and Perform channel splicing operation. It is a three-dimensional feature volume.
[0039] Furthermore, the non-destructive testing and identification method for high-strength steel welding defects based on multimodal data fusion provided by the present invention includes the following defect identification and classification methods:
[0040] Defect feature association and segmentation:
[0041] Fusion features The input to the decoder decodes the two-dimensional pixel data containing spatial geometric information to restore the three-dimensional pixel data, while simultaneously fusing features. The feature information contained therein is associated with the pixel position, and the decoder also outputs a binary segmentation mask. , The value is only 0 or 1. 1 indicates that the pixel is a defective pixel, and 0 indicates that the pixel is background. This is achieved through a binary segmentation mask. Identify defective areas;
[0042] Defect feature quantification:
[0043] For binary segmentation mask For each defective pixel connected component, the projection matrix is used. Map the defect area onto 3D point cloud data and extract all 3D points belonging to the defect. ,in This indicates the number of three-dimensional points belonging to the defect. The sequence number representing the number of 3D points of the defect is a positive integer. Project the point cloud of the defect onto the pixel coordinates of a 2D image, calculate the minimum bounding rectangle or ellipse of the defect point set, and obtain the length of the defect. and width ; Calculate all points in the defect point cloud Towards the average depth of the surrounding area maximum difference As the depth of the defect .
[0044] Defect Classification:
[0045] The feature information associated with each defective pixel connected component, and the length of the defect. ,width depth The three-dimensional spatial dimensions are matched as data pairs and input into a pre-trained fully connected neural network classifier (MLP) to automatically classify defect categories, including porosity, cracks, lack of fusion, undercut, and weld beads.
[0046] Furthermore, the non-destructive testing and identification method for high-strength steel welding defects based on multimodal data fusion provided by the present invention also includes:
[0047] Defect results output: The defect categories are labeled on the original 2D image and a defect analysis report containing metadata, inspection summary and welding defect details is output.
[0048] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0049] This invention provides a non-destructive testing and identification method for high-strength steel welding defects based on multimodal data fusion. Through welding data acquisition, data spatial alignment, feature extraction and fusion, defect identification and classification, and defect result output, it achieves non-destructive testing and identification of the three-dimensional dimensions of high-strength steel welding defects. By introducing a reference point-based precise spatial alignment method and an innovative cross-modal attention fusion network, it solves the problem of not being able to obtain the three-dimensional quantitative dimensions of defects using only two-dimensional images. Through a deep fusion and synergistic enhancement mechanism of three-dimensional geometric features and two-dimensional texture features at the intrinsic feature level, it effectively suppresses the interference of factors such as welding spatter and surface oxidation on the accuracy of defect identification. This achieves the effect of accurate identification and precise quantitative analysis of high-strength steel welding defects, solving the technical problem of high false detection and false negative rates in complex scenarios such as subtle changes in surface texture or texture interference caused by lighting, oil, and spatter in two-dimensional images. Attached Figure Description
[0050] Figure 1 This is a flowchart of a non-destructive testing and identification method for welding defects in high-strength steel based on multimodal data fusion;
[0051] Figure 2 This is a flowchart of welding data acquisition;
[0052] Figure 3 This is a flowchart of data space alignment;
[0053] Figure 4 This is a flowchart of feature extraction and fusion;
[0054] Figure 5 This is a flowchart of defect identification and classification. Detailed Implementation
[0055] The present invention will now be described in detail with reference to the accompanying drawings. The advantages and features of the present invention will become clearer from the following description. It should be noted that the drawings are all in a very simplified form and use non-precise proportions, and are only used to facilitate and clarify the illustration of the embodiments of the present invention.
[0056] Please refer to Figures 1 to 5 This invention provides a non-destructive testing and identification method for welding defects in high-strength steel based on multimodal data fusion, comprising the following steps:
[0057] Step S1, Welding data acquisition:
[0058] Two-dimensional images and three-dimensional point cloud data of the high-strength steel weld joint were acquired to form a data pair. To ensure data consistency, the two-dimensional images and three-dimensional point cloud data were acquired synchronously and from the same viewpoint, laying the foundation for subsequent data fusion processing. Specifically, this included:
[0059] Step S1.1: Set up reference points.
[0060] No fewer than four visual reference points are symmetrically set on the base material on both sides of the weld after high-strength steel welding. The reference points are high-contrast circular solid markers used for two-dimensional image localization and three-dimensional point cloud data segmentation. The circular solid markers can be black and white alternating circular coded markers.
[0061] Step S1.2: Deploy the data acquisition equipment.
[0062] A defect detection path is established, with multiple detection stations distributed along it. Acquisition devices are deployed at each detection station, ensuring that both the weld and reference point are aligned with the center of the acquisition device's field of view. The weld path can be used as the defect detection path. The acquisition devices integrate an industrial camera and a 3D scanner. The industrial camera has a resolution of at least 1200W pixels and a color depth of 24 bits, used to acquire RGB 2D images of the weld. The 3D scanner has a single-frame accuracy of ≤ ±0.05mm, ensuring precise quantification of the defect's 3D dimensions, and a point spacing of ≤ 0.1mm, used to acquire 3D point cloud data of the weld.
[0063] Step S1.3, data collection and matching.
[0064] A data acquisition command is issued to the acquisition device, and the industrial camera and 3D scanner simultaneously begin operation. The industrial camera captures a 2D image of the weld seam, while the 3D scanner acquires 3D point cloud data of the weld seam. The acquisition device correlates and stores the 2D image and 3D point cloud data to form data pairs, which serve as input for subsequent data processing. This correlation and storage is equivalent to data matching. To improve data acquisition accuracy, the weld seam surface can be cleaned according to the welding process before data acquisition.
[0065] Step S2, Data Space Alignment:
[0066] Establish a mapping relationship between pixels in a 2D image and 3D point cloud data to form a depth map corresponding to the 2D image, generating spatially aligned image-depth map data pairs. Specifically, this includes:
[0067] Step S2.1: Detection of reference points in two-dimensional image.
[0068] Image recognition techniques such as Hough circle transform or template matching are used to identify reference points in a 2D image. Subpixel-level edge extraction is performed on the identified reference points, and the geometric center pixel coordinates are calculated by ellipse fitting. ,in and These are pixel coordinate values. ,in i As a reference point, NThe integer is a positive integer, thus obtaining multiple reference points in the two-dimensional image.
[0069] Step S2.2: Detection of reference points in 3D point cloud data.
[0070] The acquired 3D point cloud data is subjected to pass-through filtering, retaining only the 3D point cloud data of the area near the weld. Then, statistical filtering or radius filtering methods are used to remove outlier noise points. For the filtered 3D point cloud data, the normal vector and curvature of each point are calculated. Since the reference point has significant curvature characteristics distinct from the weld, the point cloud data corresponding to the reference point is initially extracted through threshold segmentation. Distance-based clustering algorithms, such as Euclidean clustering, are used to group the initially extracted 3D point cloud data, with each cluster representing a reference point. Model fitting algorithms, such as random sampling consensus, are used to calculate the 3D coordinates of the center point of each cluster. This yields the position coordinates of the reference point in the three-dimensional coordinate space, where , and These are spatial coordinate values. ,in i The number of reference points, N It is a positive integer.
[0071] Steps S2.1 and S2.2 are parallel steps, and their order can be interchanged.
[0072] Step S2.3: Solve for the coordinate transformation matrix.
[0073] Using the pixel coordinates of the two-dimensional image as a reference, the three-dimensional coordinates are projected onto the two-dimensional pixel coordinates according to the coordinate transformation matrix:
[0074] (1);
[0075] In equation (1), It is a scaling factor. These are pixel coordinates. These are the three-dimensional coordinates of the point cloud data; It is the camera intrinsic parameter matrix. This can be obtained through industrial camera calibration. yes The rotation matrix, yes The translation vector, It is the camera extrinsic parameter matrix, which is to be solved.
[0076] in:
[0077] (2);
[0078] In equation (2), It is the camera intrinsic parameter matrix. and For focal length parameters, The coordinates of the main point.
[0079] in:
[0080] (3);
[0081] In equation (3), It is the camera extrinsic parameter matrix to be solved. It is the x-axis component of the translation vector. It is the y-axis component of the translation vector. It is the z-axis component of the translation vector, and r is the row and column element.
[0082] According to formula (1), the results obtained in steps S2.1 and S2.2 are... Matching pairs and the known camera intrinsic parameter matrix As input, PnP algorithms such as EPnP and IPPE are used to solve for the camera extrinsic parameter matrix. .
[0083] Step S2.4: Data space alignment.
[0084] Based on the projection matrix, the 3D point cloud data is projected onto the 2D RGB image pixel coordinate system:
[0085] (4);
[0086] In equation (4), These are the theoretical pixel coordinates corresponding to the point cloud data. This represents an equivalence relation in homogeneous coordinates, where both sides differ only by a non-zero scale factor. , It is a projection matrix. The coordinates of the point cloud data are 3D coordinates.
[0087] in:
[0088] (5);
[0089] In equation (5), It is the camera intrinsic parameter matrix. It is the camera extrinsic matrix.
[0090] All projected points are determined based on their theoretical pixel coordinates and the original... Values (i.e., depth) are resampled to generate a depth map with the exact same resolution as the 2D image. At this point, for any point in the 2D image... The corresponding depth value can be found through the depth map. This forms aligned image-depth map data pairs.
[0091] When multiple point cloud data are projected onto the same pixel coordinate, the nearest neighbor principle is used to retain the point with the smallest depth to reflect the depth value at the surface, and finally the aligned image-depth map data pair is obtained.
[0092] Step S3, Feature Extraction and Fusion:
[0093] Based on image-depth map data, feature maps and 3D feature volumes of 2D images are extracted to obtain 2D visual features and 3D feature volumes with the same number of feature channels. The 3D feature volume is used to guide the attention of the 2D visual features to generate a spatial attention weight map. This spatial attention weight map is then used to weight the 2D visual features with geometric information, and the features are fused with the 3D feature volume to obtain a fused feature containing both spatial geometric and 2D visual information. Specifically, this includes:
[0094] Step S3.1: Two-dimensional image feature extraction.
[0095] Using a 2D image from an image-depth map data pair as input, a pre-trained deep convolutional neural network (CNN) is employed as the backbone network to extract edge and texture features from the 2D image. The last fully connected layer in the CNN is removed, retaining only the convolutional and pooling layers. After the input 2D image propagates forward through the backbone network, a high-dimensional feature map is obtained at the last convolutional layer. , It is The three-dimensional tensor, in which and These represent the height and width of the feature map, respectively. It is the number of feature channels. The visual appearance information of the weld area is encoded.
[0096] Step S3.2: 3D point cloud feature extraction. This includes:
[0097] Step S3.2.1: Voxelize the 3D point cloud data. This transforms the irregular 3D point cloud into structured data suitable for convolutional network processing. Specifically, the 3D space is divided into uniform, tiny cubic meshes (i.e., voxels), with a resolution set to... ,in , and These represent the number of voxels in the depth, height, and width directions, respectively. For each non-empty voxel, the average coordinates of its interior points are calculated as a feature value, thereby transforming a sparse point cloud into a dense feature voxel mesh.
[0098] Step S3.2.2: Input the obtained feature voxel mesh into a 3D convolutional neural network (3DCNN) to extract three-dimensional geometric feature information such as concavity / convexity, curvature, and depth changes from the point cloud data. After forward propagation processing by the 3DCNN, the three-dimensional feature volume is obtained. , It is The four-dimensional tensor, in which , and These are the depth, height, and width of the feature volume, respectively. It is the number of channels in the three-dimensional feature. The spatial geometric information of the weld area is encoded. The 3D convolutional neural network (3DCNN) consists of multiple 3D convolutional layers and 3D pooling layers.
[0099] Step S3.3: Multimodal feature fusion. This includes:
[0100] Step S3.3.1: Calculate the three-dimensional feature volume Spatial dimensions and feature diagrams Alignment.
[0101] For three-dimensional feature volume Along the depth direction (i.e.) Max pooling is performed on the dimension to obtain a 3D tensor Subsequently, bilinear interpolation was used to... Perform upsampling to make its size become ; for feature maps and after upsampling pass Convolution adjusts the number of feature channels to be consistent, resulting in a size of Two-dimensional visual features of three-dimensional tensors and three-dimensional feature volume .
[0102] Step S3.3.2: Feature fusion.
[0103] Using 3D feature volume To guide two-dimensional visual features Attention generation spatial attention weight map That is, a spatial attention weight map is generated through a lightweight subnetwork. :
[0104] (6);
[0105] In equation (6), A represents the spatial attention weight map. This refers to the Sigmoid function, used to compress weights to... interval; and They are respectively The weights and biases of convolution. This represents the convolution operation. It is a three-dimensional feature volume.
[0106] Spatial attention weight map The closer the value in the value is to 1, the more significant the geometric anomaly (such as concavity, convexity, depth variation, etc.) exists at that location in three-dimensional space, and the more attention should be paid to the visual features of the corresponding location in the two-dimensional image.
[0107] Using the spatial attention weight map Two-dimensional visual features Perform geometric information weighting and combine it with three-dimensional feature volume. Feature fusion is performed to obtain fused features containing both spatial geometric information and two-dimensional visual information:
[0108] (7);
[0109] in:
[0110] (8);
[0111] In equations (7) and (8), As a feature of fusion, and They are respectively The weights and biases of convolution. This represents the convolution operation. For the weighted features, Indicates will and Perform channel splicing operation. This indicates an element-wise multiplication operation. For three-dimensional feature volume, It is a two-dimensional visual feature.
[0112] By fusing multimodal feature data, false detection of welding defects under single-modal data is effectively suppressed, and the signal weight of real defects in multimodal data is amplified, which significantly improves the accuracy of identifying complex and minute welding defects.
[0113] Step S4, Defect Identification and Classification:
[0114] The decoder restores the spatial dimensions of the fused features and associates feature information with pixel locations. It then quantizes the features of identified defective pixels, calculates the three-dimensional dimensions of the defects, and uses the feature information and three-dimensional dimensions associated with the connected components of each defective pixel as input. A pre-trained fully connected neural network classifier automatically classifies the defect categories. Specifically, this includes:
[0115] Step S4.1: Defect feature association and segmentation.
[0116] The fusion features obtained in step S3 The input is fed into the decoder. The decoder consists of multiple upsampling layers (such as transposed convolutions) and convolutional layers, which can progressively process low-resolution, high-semantic features. The spatial dimensions are restored, and the features are... The feature information contained therein (edges, textures, bumps, depth variations, etc.) is associated with pixel positions, and the decoder finally outputs a binary segmentation mask. , The value is only 0 or 1. 1 indicates that the pixel is a defect pixel (i.e., a pixel associated with defect feature information), and 0 indicates that the pixel is background. (Mask) Used to precisely locate welding defects, i.e., through a mask Identify defective areas.
[0117] Step S4.2: Defect feature quantification.
[0118] For masks For each defective pixel connected component (i.e., each independent defect instance), the projection matrix obtained in step S2 is used. mask The defective region is mapped onto 3D point cloud data, and all 3D points belonging to the defect are extracted. ,in This indicates the number of three-dimensional points belonging to the defect. The sequence number representing the number of 3D points of the defect is a positive integer. Then, the point cloud of the defect is projected onto the pixel coordinates of a 2D image using the method in step S2. The minimum bounding rectangle or ellipse of the defect point set is calculated to obtain the length of the defect. and width ; Calculate all points in the defect point cloud In the direction of depth and the average depth of the surrounding area maximum difference This transforms the visual representation of defects into quantifiable physical parameters, namely the length of the defect. ,width and depth .
[0119] Step S4.3: Defect classification.
[0120] The feature information (edges, textures, bumps, depth variations, etc.) associated with the connected components of each defective pixel, along with the physical parameters quantized in step S4.2 at that defect location, are matched as data pairs and input into a pre-trained fully connected neural network classifier (MLP) to automatically classify the defect category. The pre-trained MLP can non-linearly map "feature information (edges, textures, bumps, depth variations, etc.) + 3D spatial geometric physical parameters" to a welding defect classification library for automatic defect category classification. Defect categories include, but are not limited to, porosity, cracks, lack of fusion, undercut, and weld beads. For each independent defect instance, the MLP ultimately outputs the most probable defect category. The welding defect classification library includes information such as the true defect categories and physical size limits of the defects, and the MLP is trained based on this defect classification feature library.
[0121] Step S5, Defect Result Output: Mark the defect category on the original two-dimensional image and output a defect analysis report containing metadata, detection summary and welding defect details.
[0122] For each individual defect instance, a binary segmentation mask is used. The defect categories obtained in step S4 are labeled at the corresponding positions in the original 2D image. Simultaneously, for each independent defect instance, an integrated defect analysis report is output. The report includes metadata (inspection timestamp and operator information), an inspection summary (total number of defects, defect category statistics, maximum defect size), and a welding defect details table (containing detailed information for all defect instances). Using the labeled 2D image and the defect analysis report, defect locations can be quickly identified, and corresponding repair plans can be formulated based on the defect's physical parameters.
[0123] This invention provides a non-destructive testing and identification method for high-strength steel welding defects based on multimodal data fusion. Through welding data acquisition, data spatial alignment, feature extraction and fusion, defect identification and classification, and defect result output, it achieves non-destructive testing and identification of the three-dimensional dimensions of high-strength steel welding defects. By introducing a reference point-based precise spatial alignment method and an innovative cross-modal attention fusion network, it solves the problem of not being able to obtain the three-dimensional quantitative dimensions of defects using only two-dimensional images. Through a deep fusion and synergistic enhancement mechanism of three-dimensional geometric features and two-dimensional texture features at the intrinsic feature level, it effectively suppresses the interference of factors such as welding spatter and surface oxidation on the accuracy of defect identification, thereby achieving accurate identification and quantitative analysis of high-strength steel welding defects. This solves the technical problem of high false detection and false negative rates in complex scenarios such as subtle changes in surface texture or texture interference caused by lighting, oil, and spatter in two-dimensional images.
[0124] This invention provides a non-destructive testing and identification method for welding defects in high-strength steel based on multimodal data fusion. It achieves accurate quantification of defects and can precisely output key three-dimensional spatial geometric and physical parameters such as the length, width, and depth of defects. This provides objective and reliable direct data support for the evaluation of welding quality in high-strength steel, avoids subjective misjudgment, and has important evaluation value for ensuring the welding quality of key high-strength steel structures.
[0125] This invention provides a non-destructive testing and identification method for high-strength steel welding defects based on multimodal data fusion. By deeply fusing the rich texture and color features of two-dimensional images with the precise geometric features of three-dimensional point clouds, and achieving collaborative analysis and mutually enhanced intelligent identification of welding defects at the feature level, the deep fusion of multimodal data leverages the synergistic effect of two-dimensional images and three-dimensional point cloud data to achieve more accurate identification and quantitative analysis of high-strength steel welding defects, thereby improving the identification efficiency of high-strength steel welding defects.
[0126] This invention provides a non-destructive testing and identification method for high-strength steel welding defects based on multimodal data fusion. This method changes the existing technology, which can only identify defects through two-dimensional images or only combines two-dimensional images with three-dimensional data in a simple and superficial way, thus failing to fundamentally solve the problem of information complementarity.
[0127] This invention is not limited to the specific embodiments described above. Obviously, the embodiments described above are only a part of the embodiments of this invention, not all of them. All other embodiments obtained by those skilled in the art based on the described embodiments of this invention are within the scope of protection of this invention. Those skilled in the art can make other modifications and variations to this invention. Therefore, if these modifications and variations of this invention fall within the scope of the claims of this invention, then this invention also intends to include these modifications and variations.
Claims
1. A non-destructive testing and identification method for welding defects in high-strength steel based on multimodal data fusion, characterized in that, include: Welding data acquisition: Acquire two-dimensional images and three-dimensional point cloud data of the weld joint of high-strength steel to form data pairs; Data space Alignment: Establish the mapping relationship between the pixels of the 2D image and the 3D point cloud data to form a depth map corresponding to the 2D image, and generate spatially aligned image-depth map data pairs; Feature extraction and fusion: Based on 2D image-depth map data, feature maps and 3D feature volumes of the 2D image are extracted to obtain 2D visual features and 3D feature volumes with the same number of feature channels. Multimodal feature fusion: The 3D feature volumes are then fused together. Spatial dimensions and feature diagrams Align to obtain dimensions of Two-dimensional visual features of three-dimensional tensors and three-dimensional feature volume Using three-dimensional feature volume To guide two-dimensional visual features Attention generation spatial attention weight map That is, a spatial attention weight map is generated through a lightweight subnetwork. : (6); In equation (6), A represents the spatial attention weight map. This refers to the Sigmoid function, used to compress weights to... interval; and They are respectively The weights and biases of convolution. This represents the convolution operation. It is a three-dimensional feature volume; Using the spatial attention weight map Two-dimensional visual features Perform geometric information weighting and combine it with three-dimensional feature volume. Feature fusion is performed to obtain fused features containing both spatial geometric information and two-dimensional visual information: (7); In equation (7), As a feature of fusion, and They are respectively The weights and biases of convolution. This represents the convolution operation. For the weighted features, Indicates will and Perform channel splicing operation. It is a three-dimensional feature volume; Defect identification and classification: The decoder restores the spatial dimensions of the fused features and associates the feature information with the pixel position. The identified defective pixels are quantized to calculate the three-dimensional dimensions of the defects. The feature information and three-dimensional dimensions associated with the connected components of each defective pixel are used as input. A pre-trained fully connected neural network classifier is used to automatically classify the defect categories.
2. The method for non-destructive testing and identification of welding defects in high-strength steel based on multimodal data fusion according to claim 1, characterized in that, The method for acquiring welding data includes: Arrange reference points: Set no fewer than four visual reference points symmetrically on the base material on both sides of the weld after high-strength steel welding; Deploy acquisition equipment: Set up a defect detection path, distribute multiple detection stations along the defect detection path, and deploy acquisition equipment at the detection stations, wherein the acquisition equipment integrates an industrial camera and a 3D scanner. Data acquisition and matching: Two-dimensional images of the weld are acquired using an industrial camera, and three-dimensional point cloud data of the weld are acquired using a 3D scanner. The acquisition equipment then associates and stores the two-dimensional images and three-dimensional point cloud data to form data pairs.
3. The method for non-destructive testing and identification of welding defects in high-strength steel based on multimodal data fusion according to claim 1, characterized in that, The data space alignment method includes: 2D Image Reference Point Detection: Reference points in a 2D image are identified using image recognition technology. Subpixel-level edge extraction is performed on the identified reference points, and the geometric center pixel coordinates are calculated through ellipse fitting. ,in and These are pixel coordinate values. ,in i The number of reference points, N A positive integer, which represents the coordinates of a reference point in the two-dimensional image as the base point; 3D point cloud data reference point detection: The acquired 3D point cloud data is subjected to pass-through filtering to retain the 3D point cloud data near the weld seam. The data near the weld seam is then filtered to remove outlier noise points. For the filtered 3D point cloud data, the normal vector and curvature of each point are calculated. Threshold segmentation is used to initially extract the point cloud data corresponding to the reference points. A distance-based clustering algorithm is then used to group the initially extracted 3D point cloud data, with each cluster representing a reference point. The 3D coordinates of the center point of each cluster are calculated. This yields the three-dimensional coordinates of the reference point's position, where , and These are spatial coordinate values. ,in i The number of reference points, N It is a positive integer; Solve for the coordinate transformation matrix: Using the pixel coordinates of the two-dimensional image as a reference, the three-dimensional coordinates are projected onto the two-dimensional pixel coordinates according to the coordinate transformation matrix: (1); In equation (1), It is a scaling factor. These are pixel coordinates. These are the three-dimensional coordinates of the point cloud data; It is the camera intrinsic parameter matrix. yes The rotation matrix, yes The translation vector, It is the camera extrinsic parameter matrix; According to formula (1), the matching point pairs and the known camera intrinsic parameter matrix As input, solve for the camera extrinsic matrix. ; Data spatial alignment: Based on the projection matrix, the 3D point cloud data is projected onto the 2D image pixel coordinate system: (4); In equation (4), These are the theoretical pixel coordinates corresponding to the point cloud data. Represents the equivalence relation in homogeneous coordinates. It is a projection matrix. The three-dimensional coordinates of the point cloud data; All projected points are determined based on their theoretical pixel coordinates and the original... Value resampling generates a depth map with the same resolution as the two-dimensional image, forming an aligned image-depth map data pair.
4. The method for non-destructive testing and identification of welding defects in high-strength steel based on multimodal data fusion according to claim 3, characterized in that, The feature extraction and fusion methods include: Two-dimensional image feature extraction: Using a 2D image from an image-depth map data pair as input, a pre-trained deep convolutional neural network (CNN) is used as the backbone to extract feature information from the 2D image, including edges and textures. After the input 2D image propagates forward through the backbone network, a high-dimensional feature map is obtained in the last convolutional layer. , It is The three-dimensional tensor, in which and These represent the height and width of the feature map, respectively. It is the number of feature channels. The visual appearance information of the weld area is encoded; 3D point cloud feature extraction: Voxelization of 3D point cloud data is performed to form a feature voxel mesh; The feature voxel mesh is input into a 3D convolutional neural network (3DCNN) to extract the 3D geometric feature information of points in the point cloud data. This 3D geometric feature information includes convexity, curvature, and depth variations. After forward propagation processing by the 3DCNN, the 3D feature volume is obtained. , It is The four-dimensional tensor, in which , and These are the depth, height, and width of the feature volume, respectively. It is the number of channels for the three-dimensional feature.
5. The method for non-destructive testing and identification of welding defects in high-strength steel based on multimodal data fusion according to claim 4, characterized in that, The defect identification and classification method includes: Defect feature association and segmentation: Fusion features The input to the decoder decodes the two-dimensional pixel data containing spatial geometric information to restore the three-dimensional pixel data, while simultaneously fusing features. The feature information contained therein is associated with the pixel position, and the decoder also outputs a binary segmentation mask. , The value is only 0 or 1. 1 indicates that the pixel is a defective pixel, and 0 indicates that the pixel is background. This is achieved through a binary segmentation mask. Identify defective areas; Defect feature quantification: For binary segmentation mask For each defective pixel connected component, the projection matrix is used. Map the defect area onto 3D point cloud data and extract all 3D points belonging to the defect. ,in This indicates the number of three-dimensional points belonging to the defect. The sequence number representing the number of 3D points of the defect is a positive integer. Project the point cloud of the defect onto the pixel coordinates of a 2D image, calculate the minimum bounding rectangle or ellipse of the defect point set, and obtain the length of the defect. and width ; Calculate all points in the defect point cloud Towards the average depth of the surrounding area maximum difference As the depth of the defect ; Defect Classification: The feature information associated with each defective pixel connected component, and the length of the defect. ,width depth The three-dimensional spatial dimensions are matched as data pairs and input into a pre-trained fully connected neural network classifier (MLP) to automatically classify defect categories, including porosity, cracks, lack of fusion, undercut, and weld beads.
6. The method for non-destructive testing and identification of welding defects in high-strength steel based on multimodal data fusion according to claim 1, characterized in that, Also includes: Defect results output: The defect categories are labeled on the original 2D image and a defect analysis report containing metadata, inspection summary and welding defect details is output.
Citation Information
Patent Citations
Metal weld defect detection method based on multi-modal fusion and defect detection network framework
CN118608479A
Defect detection method for semiconductor packaging material based on deep learning
CN120525859A