Weld defect intelligent detection method based on high-definition image deep learning

By building a weld defect image database and multi-stage fine-tuning of the GLIP model, combined with multimodal data fusion technology, the problem that traditional weld inspection methods cannot identify internal defects is solved, and efficient and accurate weld defect detection is achieved.

CN120707537APending Publication Date: 2025-09-26GUANGXI UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510856252.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing weld defect detection methods cannot effectively identify internal defects in welds and lack the ability to identify complex defects. Traditional non-destructive testing methods rely on single modal data and find it difficult to achieve multi-dimensional quantitative analysis.

Method used

A high-quality weld defect image database is constructed, the GLIP model is used for two-stage fine-tuning, a multi-layer perceptron module is introduced for multimodal feature alignment, and cross-modal feature fusion is performed by combining image and ultrasonic data to achieve joint detection of internal and external defects in welds.

Benefits of technology

It significantly improves the automation, accuracy and robustness of weld defect detection, can simultaneously identify weld surface and internal defects, and enhances adaptability to complex working conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention provides a weld defect intelligent detection method based on high-definition image deep learning, which comprises the following steps: S1, collecting a high-definition weld image, manually marking the selected image, and constructing a weld defect image database; s2, performing first fine tuning on the pre-training weight GLIP-L by adopting a GLIP model; s3, a self-made welding test piece is made, a high-definition camera and ultrasonic detection equipment are used for collecting a weld joint surface high-definition image of the self-made welding test piece and ultrasonic echo characteristic data of internal defects, and a high-definition weld joint image and ultrasonic multi-mode database is established; s4, introducing a multi-layer perceptron module into the GLIP model, constructing an image and ultrasonic multi-modal feature alignment mechanism, and performing second fine tuning to obtain a trained GLIP model; and S5, performing defect detection on the welding seam by using the trained GLIP model, and outputting a detection result. According to the invention, through multi-modal data fusion, the accuracy and robustness of weld defect detection are improved, and the method is suitable for automatic and accurate industrial detection scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of computer vision, deep learning, data science, and engineering applications, and specifically relates to an intelligent detection method for weld defects based on high-definition image deep learning. Background Art

[0002] Welding technology has been widely used in all fields, from heavy industry to precision manufacturing, but its quality defects can cause major safety hazards, and there is an urgent need to establish an efficient and accurate weld defect detection system. Currently, weld defect detection mainly relies on destructive testing and manual visual inspection, which are difficult to meet the needs of modern manufacturing for efficient and accurate testing. At the same time, non-destructive testing technology is widely used in the industrial field because it can accurately identify internal and surface defects without destroying the object being tested. However, traditional non-destructive testing methods still have problems such as high dependence on operator experience, low data processing efficiency, and difficulty in achieving multi-dimensional quantitative analysis when dealing with complex structures or minor defects.

[0003] In recent years, with the rapid development of deep learning and computer vision technologies, neural network-based object detection technology has matured, providing a novel solution for nondestructive testing (NDT). These technologies offer advantages such as high detection efficiency, high automation, and high recognition accuracy. Their application in infrastructure can effectively overcome the shortcomings of traditional NDT methods. GLIP is a multi-task unified model based on visual language understanding. By deeply fusing visual features with semantic descriptions, it achieves efficient synergy between object detection and semantic understanding. The model employs a dual-encoder architecture, combining visual features extracted by an image encoder with semantic descriptions generated by a text encoder. A cross-attention mechanism is used to establish a cross-modal feature interaction channel, enabling precise localization and semantic association of target regions. This mechanism enables the GLIP model to perform zero-shot object detection directly from textual cues, breaking the limitation of traditional detection models that rely on fixed category annotations. To address the engineering challenge of significant scale variation in weld defects, GLIP innovatively introduces a multi-scale feature fusion mechanism. Specifically, the GLIP model retains high-resolution details in low-level feature maps to capture microscopic defects, while leveraging high-level feature maps to extract abstract semantic features for macroscopic defect recognition. Cross-scale feature interaction enhances information flow between different layers. This hierarchical and progressive feature processing strategy enables the GLIP model to simultaneously achieve micron-level precision detection and global perception of complex defect clusters within a unified framework. However, current image-based weld defect detection methods have the limitation of only being able to identify surface defects but not internal defects in welds. Summary of the Invention

[0004] In response to the problems existing in the above-mentioned prior art, the present invention provides an intelligent weld defect detection method based on high-definition image deep learning. The purpose is to achieve efficient joint detection of internal and external defects in welds by constructing a high-quality weld defect image database, fine-tuning the GLIP model in two stages, establishing a multimodal database and a semi-automatic annotation process, and optimizing the performance of the GLIP model through cross-modal feature alignment. The reliability of the GLIP model is verified by combining multimodal data, thereby solving the problem that traditional weld detection relies on single-modal data and lacks the ability to identify complex defects. The method has high detection accuracy and robustness.

[0005] In order to achieve the above object, the specific scheme of the present invention is as follows:

[0006] S1, collect high-definition weld images and randomly sample them, and manually annotate the selected randomly sampled images. The annotation content includes defect type and spatial coordinate information to build a weld defect image database;

[0007] S2, based on the weld defect image database built in S1, uses the GLIP model to perform the first fine-tuning of the pre-trained weights GLIP-L;

[0008] S3: Prepare a self-made welding specimen and use a high-definition camera and ultrasonic testing equipment to collect high-definition images of the weld surface and ultrasonic echo characteristic data of internal defects of the self-made welding specimen, and establish a high-definition weld image and ultrasonic multimodal database;

[0009] S4, based on the high-definition weld images and ultrasonic multimodal database built in S3, introduces a multi-layer perceptron module into the GLIP model, builds an image and ultrasonic multimodal feature alignment mechanism, and performs a second fine-tuning to obtain the trained GLIP model;

[0010] S5, using the GLIP model trained in step S4 to perform defect detection on the weld and output the detection results.

[0011] Furthermore, the method further comprises the following steps:

[0012] S51, based on S3 high-definition weld images and ultrasonic multimodal database, builds a multimodal weld defect test set;

[0013] In step S52, the performance of the GLIP model trained in step S4 is evaluated using mAP, IoU, and recall metrics to verify the performance of defect detection in step S5.

[0014] Furthermore, the collection of high-definition weld images in step S1 is to collect 37,150 high-definition weld images from a network resource library, and randomly sample and screen out 5,930 images to construct a training set for the weld defect image database; the manual annotation is to use Labelme annotation software to systematically manually annotate the selected images, and the defect types include spatter, cracks, pores, inclusions, dents, scratches, incomplete penetration and burrs.

[0015] Furthermore, the first fine-tuning of the pre-trained weight GLIP-L in step S2 includes the following steps:

[0016] S21, load the pre-trained weight GLIP-L and specify the storage path of the pre-trained weight file of the GLIP model;

[0017] S22, set fine-tuning structure parameters: selectively freeze the parameters of the underlying visual encoder to retain the general image features learned in the pre-training phase; selectively unfreeze the parameters of the high-level semantic interaction module to enable the GLIP model to learn specific features related to weld defect detection;

[0018] S23, applies a hierarchical freezing strategy to update only the parameters of the high-level semantic interaction module while keeping the parameters of the underlying visual encoder unchanged, in order to optimize the GLIP model's ability to recognize weld features;

[0019] S24: Build a dedicated configuration system for image annotation tasks. The dedicated configuration system for image annotation tasks includes:

[0020] Data path configuration, used to clearly mark the JSON file and image storage path output by Labelme annotation software;

[0021] Defect semantic definition: establishing a text-image mapping relationship through defect descriptive prompt words;

[0022] Task mode setting, used to specify the detection task type, image annotation task and output COCO format;

[0023] Enhancement strategy setting, used for brightness, contrast adjustment and horizontal flip data enhancement;

[0024] Distributed training settings, used to activate dual-GPU parallel training mode and accelerate GLIP model convergence;

[0025] Fine-tune strategy selection and use the Language_Prompt v4 strategy to optimize the text prompt mechanism;

[0026] S25 , performing a first fine-tuning based on the pre-trained weights and parameter settings loaded in steps S21 , S22 , S23 , and S24 .

[0027] Furthermore, the preparation of the self-made welding specimen in step S3 includes the following steps:

[0028] For S31, Q345 steel and S32168 stainless steel were selected as welding materials, and manual arc welding and gas shielded welding processes were used to simulate actual production scenarios to prepare welding specimens;

[0029] S32, artificially introduces defects such as cracks, pores, spatter and dents into the weld specimen to simulate the defects in the actual welding process;

[0030] S33, V-shaped or U-shaped grooves are processed by CNC machine tools, and combined with laser cleaning to ensure the surface cleanliness of the welding specimens; the assembly accuracy of the welding specimens is guaranteed by fixture positioning and laser calibration, forming a structural form covering butt welds and fillet welds. The thickness of the welding specimens ranges from 12-20mm.

[0031] Furthermore, the secondary fine-tuning in step S4 includes the following steps:

[0032] S41, by referencing the multi-layer perceptron module in the GLIP model, the appearance characteristics of the weld image and the time domain signal characteristics of the ultrasonic signal are cross-modally mapped;

[0033] S42, unfreeze all frozen layers, adjust the multilayer perceptron module parameters, including the number of channels and activation function, and adjust the training strategy, including the learning rate decay step size and gradient clipping threshold;

[0034] S43: Build a fine-tuning configuration file dedicated to the target detection task, update the dataset path, switch the task type to target detection, and use the Full Fine-tuning fine-tuning strategy.

[0035] S44. Based on steps S41, S42 and S43, the GLIP model is fine-tuned for the second time in combination with a multimodal joint loss function. The joint loss function includes classification loss, regression loss and cross-modal consistency loss. The classification loss uses Focal Loss to alleviate the category imbalance problem, the regression loss uses CIOU Loss to accurately evaluate the positioning error of the target box, and the cross-modal consistency loss uses contrastive learning to constrain the multimodal features to be aligned in the shared semantic space.

[0036] Advantages of the present invention

[0037] 1. Efficiency of a Dual-Task Framework: This paper employs two-stage fine-tuning to construct a dual-task framework for image annotation and object detection, significantly improving the efficiency of the detection process compared to traditional single-task frameworks. Furthermore, the GLIP model, fine-tuned using the Language_Prompt V4 strategy, rapidly completes the automatic annotation of unknown weld images, achieving simultaneous output of defect bounding boxes and text descriptions, significantly improving annotation efficiency and surpassing traditional manual annotation methods.

[0038] 2. Cross-modal feature alignment: The GLIP model innovatively introduces a multi-layer perceptron (MLP) module, and by constructing a multi-modal feature alignment mechanism, the feature space of ultrasonic data and image data is deeply fused. During the training phase, the GLIP model jointly models ultrasonic signals and image data, enabling the network to synchronously learn the implicit association between apparent defects and internal defects in welds; while in the inference phase, only monocular images need to be input to achieve joint detection of internal and external defects. The present invention focuses on the two key fine-tuning processes of GLIP: the first fine-tuning uses a small amount of manually labeled data to achieve automatic labeling and improve labeling efficiency; the second fine-tuning completes the feature alignment of multi-modal data through the innovative design of the MLP layer, thereby constructing an intelligent detection framework that can automatically label and fuse internal and external defect features. This solves the problem that traditional methods cannot identify internal defects in welds, and provides new ideas for the comprehensive analysis of weld defects.

[0039] 3. End-to-end joint defect detection: Utilizing a full-fine tuning strategy and a distributed dual-GPU training mechanism, this invention achieves joint detection of internal and surface defects in welds. The GLIP model, fed with a high-definition weld image, simultaneously outputs detection results for both surface defects (e.g., cracks and pores) and internal defects (e.g., lack of fusion and slag inclusions). Compared to traditional image-based object detection methods, this invention improves detection comprehensiveness and reliability.

[0040] 4. Enhanced robustness: Through multimodal feature alignment and joint modeling, this paper leverages the complementary nature of ultrasonic and image data to enhance the robustness of the GLIP model to complex working conditions (such as illumination changes and surface noise). This particularly improves the efficiency of internal weld defect recognition, moving from low efficiency to high efficiency.

[0041] In summary, the present invention significantly improves the automation, accuracy, and robustness of weld defect detection through multimodal data fusion and deep learning technology, providing an efficient and reliable solution for weld quality control in industrial applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 This is a flow chart of a method for intelligent detection of weld defects based on high-definition image deep learning according to the present invention.

[0043] Figure 2 for Figure 1 The architecture diagram of the multilayer perceptron module introduced in the GLIP model in

[15] .

[0044] Figure 3 The sample image set of weld defects used for training and validating the GLIP model is shown.

[0045] Figure 4 The changing curves of loss function and accuracy during training and validating the GLIP model are shown.

[0046] Figure 5 for Figure 1 Schematic diagram of the structure of the homemade welding specimen.

[0047] Figure 6 The figure shows the results of weld defect detection using the trained GLIP model. DETAILED DESCRIPTION

[0048] The present invention will be further explained and illustrated below in conjunction with the accompanying drawings and specific embodiments. It should be noted that this specific embodiment is not intended to limit the scope of rights of the present invention.

[0049] like Figures 1 to 6 As shown, this specific embodiment provides a method for intelligent detection of weld defects based on high-definition image deep learning, including the following steps:

[0050] S1, collect high-definition weld images. The high-definition weld images are collected from a network resource library. 5930 images are randomly sampled and selected to construct a training set for the weld defect image database. The selected images are manually annotated to construct the weld defect image database. The manual annotation is a systematic manual annotation of the selected images using Labelme annotation software. The manual annotation content covers defect types and spatial coordinate information such as cracks, pores, spatter, and dents. The annotation content includes defect type and spatial coordinate information. The defect types include spatter, cracks, pores, inclusions, dents, scratches, incomplete penetration, and burrs.

[0051] Specifically, this embodiment is based on the existing research progress of deep learning in the field of weld defect detection and public dataset resources. In the first fine-tuning stage, public network data is used to construct a benchmark dataset.

[0052] SS1 collects high-definition weld images from an internet resource library and randomly selects 5,930 images as a labeled training set, covering eight typical defects: spatters, cracks, porosity, inclusions, undercuts, scratches, incomplete penetration, and burrs. Figure 3 As shown, it provides standardized samples for subsequent annotation tasks.

[0053] SS2 uses the Python-based image annotation tool Labelme to complete the annotation of the first fine-tuning data. The operation process is as follows:

[0054] 1. Start the Labelme annotation software and use the "Open Dir" function to load the folder where the image to be annotated is located;

[0055] 2. Select the target image in the "File List" and use the "Create Polygons" tool to draw a polygonal selection around the defect area. Simultaneously complete the defect category name (such as adj, int, etc.). For folders containing multiple images, you can use the "Next Image" or "Prev Image" button to switch to the next or previous image and repeat the above annotation process.

[0056] 3. After completing the annotation, click the "Save" button to save the annotation results in JSON format to the specified folder. The JSON content includes the shape of the target detection box, the coordinates of all points that make up the target detection box, the label name, the image path, and the image pixels.

[0057] S2, based on the weld defect image database constructed in S1, consists of 5,930 annotated high-definition weld images. A stratified sampling strategy was used to construct the dataset: 4,151 images were randomly selected as the training set (70%), 890 as the test set (15%), and 889 as the validation set (15%). Due to limited data and the need to optimize text labels, the Language_Prompt v4 strategy was selected for the first fine-tuning of the pre-trained weights of the GLIP model, GLIP-L. In this step, a dedicated object detection configuration framework was constructed, consisting of three core configuration items: dataset path definition, detection task mode selection, and learning strategy configuration, providing a standardized parameter interface for the second fine-tuning. The dataset path definition specifies the storage paths for the training, validation, and test datasets. The detection task mode selection supports both image annotation and object detection modalities, selected based on task requirements. The learning strategy configuration includes hyperparameters such as the optimizer type and training rounds, providing the necessary learning strategy for GLIP model training. Dynamic text prompts enhance the alignment of text and visual features, resolving semantic bias in small sample scenarios. Focusing on improving the GLIP model's ability to perceive weld features, the model is equipped with the ability to automatically annotate unknown weld images and output the three elements of defect location, category, and attributes.

[0058] The pre-trained weight GLIP-L fine-tuning includes the following steps:

[0059] S21, load the pre-trained weight GLIP-L, use the Language_Prompt V4 strategy in the configs / glip_Swin_L.yaml configuration file, and specify the storage path of the pre-trained weight file of the GLIP model;

[0060] S22, set fine-tuning structure parameters: selectively freeze the parameters of the underlying visual encoder to retain the general image features learned in the pre-training phase; selectively unfreeze the parameters of the high-level semantic interaction module to enable the GLIP model to learn specific features related to weld defect detection; optimize the text prompt mechanism through the Language_Prompt v4 strategy, adjust the learning rate and number of training iterations, as shown in Table 1:

[0061] Table 1 GLIP model fine-tuning parameters

[0062]

[0063] S23 applies a hierarchical freezing strategy to dynamically control the trainable state of the language backbone network according to task requirements, updating only the parameters of the high-level semantic interaction module while keeping the parameters of the underlying visual encoder unchanged, so as to optimize the GLIP model's ability to recognize weld features and ensure the flexibility of aligning text and image features.

[0064] S24, constructing a dedicated configuration system for image annotation tasks to provide a standardized parameter framework for the GLIP model optimization in step S23. The dedicated configuration system for image annotation tasks includes:

[0065] Data path configuration, used to clearly mark the JSON file and image storage path output by Labelme annotation software;

[0066] Defect semantic definition, establishes the mapping relationship between text and image through the defect descriptive prompt words shown in Table 2;

[0067] Task mode setting, used to specify the detection task type, image annotation task and output COCO format;

[0068] Enhancement strategy setting, used for brightness and contrast adjustment and horizontal flip data augmentation to improve the generalization ability of the GLIP model;

[0069] Distributed training settings, used to activate dual-GPU (2×NVIDIA 3090Ti) parallel training mode to accelerate GLIP model convergence;

[0070] Fine-tune strategy selection and use the Language_Prompt v4 strategy to optimize the text prompt mechanism.

[0071] S25 employs the Language_Prompt v4 fine-tuning strategy, leveraging dynamic language prompts and a cross-modal alignment mechanism to adapt the GLIP model for high-precision labeling of industrial weld images. This approach leverages domain-specific multimodal prompt templates (e.g., "weld region: {location}, defect type: {crack / pore}") to generate descriptive prompt words strongly associated with local image features (e.g., cracks, pores) through dynamic semantic mapping. At the cross-modal interaction layer, a multi-head attention mechanism fuses the visual and textual features of high-definition weld images, enhancing the GLIP model's fine-grained perception of subtle defects.

[0072] S26, constructs high-quality weld images and annotation data through S1, and simultaneously defines data augmentation strategies in the annotation task configuration file to perform super-resolution reconstruction and local contrast enhancement on the weld images to improve the clarity of key areas;

[0073] S27 freezes the underlying visual encoder weights of the GLIP backbone network during training, fine-tuning only the high-level cross-modal interaction layer and dynamic cue generation module. A two-stage training strategy is employed. In the first stage, contrastive learning (InfoNCE loss, weight 0.6) is used to enhance cross-modal semantic alignment. In the second stage, a GIoU bounding box regression loss (weight 0.3) and a classification loss (weight 0.1) are introduced. Combined with a curriculum learning strategy, task weights are dynamically adjusted to achieve joint optimization of defect location and type.

[0074] S28, in the inference stage, after inputting an unknown weld image, the GLIP model generates candidate semantic labels based on predefined prompt templates, decodes through visual and language feature alignment, outputs defect location and type predictions, uses non-maximum suppression (NMS) to eliminate redundant predictions, and performs logical verification in combination with the welding process knowledge base, and finally generates high-confidence annotation results (such as JSON format).

[0075] Table 2 Descriptive words for defects

[0076]

[0077] S3. Based on industry standards such as the Steel Structure Welding Code, various welding processes were used to prepare custom weld specimens containing typical defect types such as cracks, pores, spatter, and dents, providing standardized samples for subsequent data collection. Given that existing public weld defect datasets generally lack ultrasonic signal support, to achieve joint modeling of image and ultrasonic data, this embodiment uses a high-definition camera and ultrasonic testing equipment to simultaneously collect high-definition images of the weld surface and ultrasonic echo signature data of internal defects from the custom weld specimens, thereby establishing a high-definition weld image and ultrasonic multimodal database. The high-definition camera is combined with a ring-shaped LED light source to capture high-definition images of the weld surface. The ultrasonic testing equipment utilizes an Olympus Pharos 2.8 phased array ultrasonic system with a 5MHz high-frequency probe. The probe contacts the weld surface via a water-based or glycerin coupling agent. A six-axis robotic arm is used to move axially and transversely along the weld with a precision of 0.1mm. FPGA hard triggering is used to synchronously collect A-scan waveforms (i.e., time-domain signals) and B / C-scan images (i.e., two-dimensional / three-dimensional defect distributions), obtaining high-precision echo signatures of internal weld defects. Image-ultrasonic multimodal database of defects such as internal and external cracks, pores, sink marks, and lack of fusion in the final structure weld.

[0078] The preparation of the self-made welding specimen comprises the following steps:

[0079] S31, according to the Steel Structure Welding Code (GB 50661-2011) and other industrial standards, selected Q345 steel and S32168 stainless steel as welding materials, and used manual arc welding and gas shielded welding processes to simulate actual production scenarios to prepare welding specimens;

[0080] S32, artificially introduces defects such as cracks, pores, spatter and dents into the weld specimen to simulate the defects in the actual welding process;

[0081] S33, V-shaped or U-shaped grooves are processed by CNC machine tools, and laser cleaning is combined to ensure the surface cleanliness of the welding specimens; the assembly accuracy of the welding specimens is ensured by fixture positioning and laser calibration, forming a structural form covering butt welds and fillet welds, such as Figure 5 As shown, the thickness of the welding specimens ranges from 12 to 20 mm.

[0082] To overcome the limitations of traditional image detection methods for identifying internal weld defects, S4 introduced a multi-layer perceptron module into the GLIP model based on the high-definition weld images and ultrasonic multimodal database constructed in S3. This module established an alignment mechanism for image and ultrasonic multimodal features and performed a second fine-tuning to obtain the fully trained GLIP model.

[0083] The second fine-tuning includes the following steps:

[0084] S41 embeds an MLP module within the GLIP model. Based on the high-definition weld images and ultrasonic multimodal database constructed in S3, it performs cross-modal mapping between the surface features of the weld images and the time-domain signal characteristics of the ultrasonic signals. This establishes a cross-modal alignment mechanism between the surface features of defects and their internal structural features. This provides a feature fusion channel for the GLIP model to simultaneously perceive surface morphology and internal defects, and enables the GLIP model to simultaneously learn the implicit associations between surface and internal defects in the second fine-tuning phase. This phase adopts a full fine-tuning strategy, inputting training samples of high-definition weld images and ultrasonic multimodal data. This end-to-end joint modeling improves defect location accuracy and detection quality.

[0085] S42, make key adjustments to the configs / glip_Swin_L.yaml configuration file: unfreeze all frozen layers to enhance the trainability of the GLIP model for multimodal features; adjust the multilayer perceptron module parameters to enhance the cross-modal feature alignment capability; the multilayer perceptron module parameters include the number of channels and activation function, and adjust the training strategy, including the learning rate decay step size and gradient clipping threshold; ensure the stable convergence of the GLIP model under multimodal data.

[0086] S43, builds a fine-tuning configuration file dedicated to the target detection task. Its core parameter settings are highly consistent with the configuration file for the image annotation task, and only the following key parameters are adjusted in a targeted manner: 1) Dataset path update, pointing the annotation data storage path to the local directory containing weld image-ultrasonic multimodal data; 2) Task type switching, changing the image annotation task (grounding) to target detection (detection); 3) Training strategy optimization, changing the Language_Prompt v4 strategy to the Full Fine-tuning strategy, unfreezing all parameters of the GLIP model, and synchronously optimizing the multimodal feature alignment capability through gradient backpropagation.

[0087] The S44 MLP module is used to align images and ultrasonic data. It consists of two key stages: first, the high-definition weld image and ultrasonic time-domain signal features are used as multimodal input to extract visual features and ultrasonic features respectively; then, the two types of features are cross-modally fused through the newly added MLP layer to learn the correlation between the image and the ultrasonic signal.

[0088] During the training phase, Figure 4 As shown, based on steps S41, S42, and S43, the GLIP model is fine-tuned for the second time in combination with the multimodal joint loss function to improve the performance of the GLIP model, as follows:

[0089] 1) Focal Loss is used in classification loss to alleviate the problem of category imbalance, such as the low proportion of defect samples, and the category-aware attention mechanism is used to enhance the learning ability of key defect categories;

[0090] 2) CIOU Loss is used as the regression loss. By combining the intersection-over-union ratio, center point distance, and aspect ratio penalty terms, it accurately evaluates the positioning error of the target box and improves the robustness of defect boundary prediction.

[0091] 3) A cross-modal consistency loss employs contrastive learning to align multimodal features in a shared semantic space, ensuring consistent representation of the same defect between high-definition weld images and ultrasonic time-domain signals. The entire GLIP model is trained end-to-end based on multimodal weld defect data, and all parameters are simultaneously optimized through backpropagation, enabling GLIP to achieve high-precision defect detection capabilities based on multimodal data.

[0092] S5, use the trained GLIP model to detect defects in the weld and output the detection results, such as Figure 6 shown.

[0093] S51, based on S3 high-definition weld images and ultrasonic multimodal database, builds a multimodal weld defect test set.

[0094] A systematic evaluation is conducted on the trained GLIP model. The specific steps are as follows: First, randomly select no less than 2,000 groups of data samples that have not participated in the training from the database to ensure that these samples can fully cover all types of welding defects and their complexity. Subsequently, the image data are automatically annotated using the GLIP model that has been fine-tuned once in step S2. The GLIP model will output the weld defect type and the corresponding spatial coordinate information in each image, thereby constructing a set of test sets containing automatic annotation results, in which each sample is associated with the weld surface defect prediction result generated by the GLIP model and the corresponding ultrasonic feature description. Next, the test set is input into the GLIP model that has been fine-tuned for the second time in step S4. With the help of the previously established image and ultrasonic multimodal feature alignment mechanism, the effective fusion of the two modal information is achieved, thereby conducting a comprehensive and systematic performance evaluation of the trained GLIP model.

[0095] In step S52, the performance of the GLIP model trained in step S4 was evaluated using mAP, IoU, and recall metrics. Quantitatively, mAP (comprehensive detection accuracy), IoU (localization accuracy), and recall (missed detection rate) were used. Qualitative analysis was performed by visualizing defect distribution characteristics through defect location heatmaps and false detection / missed detection statistics. Ultrasonic echo feature data was introduced as an auxiliary verification tool to establish a multimodal evaluation system for "appearance-internal" defect detection. Multidimensional comparative experiments were designed to compare the improved GLIP model with the original GLIP, YOLOv8, and traditional visual inspection methods, verifying the effectiveness of the MLP layer in the GLIP model's cross-modal feature alignment mechanism. Comparisons were made with traditional methods and single-modal inspection results, focusing on the detection accuracy of the two-stage fine-tuned GLIP model for complex defects such as small cracks and deep pores, verifying the reliability and robustness of multimodal fusion in defect assessment. Detailed results are shown in Table 3. Combined with the actual needs of industrial scenarios, the robustness of the improved GLIP model in terms of complex background interference, small sample defect recognition, and inference speed is analyzed. The contribution of the MLP layer to improving the performance of the GLIP model is quantified through an ablation study. The specific results are shown in Table 4, which provides a theoretical basis and performance guarantee for the industrial implementation of weld defect detection.

[0096] Through the description of the above specific embodiments, the present invention provides an intelligent detection method for weld defects based on image and ultrasonic multimodal architecture. The method constructs a high-quality weld defect image database, fine-tunes the GLIP model in two stages, establishes a multimodal database and a semi-automatic annotation process, optimizes the performance of the GLIP model through cross-modal feature alignment, and verifies the reliability of the GLIP model in combination with multimodal data. This method realizes efficient joint detection of internal and external defects in welds, solves the problem that traditional weld detection relies on single modal data and has insufficient ability to identify complex defects, and has high detection accuracy and robustness.

[0097] Table 3 Comparison of multiple model results

[0098]

[0099] Table 4 Comparison of ablation experiment results

[0100]

Claims

1. A high-definition image deep learning intelligent detection method for weld defects, characterized by: The following steps are involved: S1, collect high-definition weld images and randomly sample them, and manually annotate the selected randomly sampled images. The annotation content includes defect type and spatial coordinate information to build a weld defect image database; S2, based on the weld defect image database built in S1, uses the GLIP model to perform the first fine-tuning of the pre-trained weights GLIP-L; S3: Prepare a self-made welding specimen and use a high-definition camera and ultrasonic testing equipment to collect high-definition images of the weld surface and ultrasonic echo characteristic data of internal defects of the self-made welding specimen, and establish a high-definition weld image and ultrasonic multimodal database; S4, based on the high-definition weld images and ultrasonic multimodal database built in S3, introduces a multi-layer perceptron module into the GLIP model, builds an image and ultrasonic multimodal feature alignment mechanism, and performs a second fine-tuning to obtain the trained GLIP model; S5, using the GLIP model trained in step S4 to perform defect detection on the weld and output the detection results.

2. The method according to claim 1, characterized in that The following steps are also included: S51, based on S3 high-definition weld images and ultrasonic multimodal database, builds a multimodal weld defect test set; In step S52, the performance of the GLIP model trained in step S4 is evaluated using mAP, IoU, and recall metrics to verify the performance of defect detection in step S5.

3. The method according to claim 1, characterized in that The collection of high-definition weld images in step S1 is to collect 37,150 high-definition weld images from a network resource library, and randomly sample and screen 5,930 images to construct a training set for the weld defect image database; the manual annotation is to use Labelme annotation software to systematically manually annotate the selected images, and the defect types include spatter, cracks, pores, inclusions, dents, scratches, incomplete penetration and burrs.

4. The method according to claim 1, wherein The first fine-tuning of the pre-trained weights GLIP-L in step S2 includes the following steps: S21, load the pre-trained weight GLIP-L and specify the storage path of the pre-trained weight file of the GLIP model; S22, set fine-tuning structure parameters: selectively freeze the parameters of the underlying visual encoder to retain the general image features learned in the pre-training phase; selectively unfreeze the parameters of the high-level semantic interaction module to enable the GLIP model to learn specific features related to weld defect detection; S23, applies a hierarchical freezing strategy to update only the parameters of the high-level semantic interaction module while keeping the parameters of the underlying visual encoder unchanged, in order to optimize the GLIP model's ability to recognize weld features; S24: Build a dedicated configuration system for image annotation tasks. The dedicated configuration system for image annotation tasks includes: Data path configuration, used to clearly mark the JSON file and image storage path output by Labelme annotation software; Defect semantic definition: establishing a text-image mapping relationship through defect descriptive prompt words; Task mode setting, used to specify the detection task type, image annotation task and output COCO format; Enhancement strategy setting, used for brightness, contrast adjustment and horizontal flip data enhancement; Distributed training settings, used to activate dual-GPU parallel training mode and accelerate GLIP model convergence; Fine-tune strategy selection and use the Language_Prompt v4 strategy to optimize the text prompt mechanism; S25 , performing a first fine-tuning based on the pre-trained weights and parameter settings loaded in steps S21 , S22 , S23 , and S24 .

5. The method according to claim 1, wherein The preparation of the self-made welding specimen in step S3 includes the following steps: For S31, Q345 steel and S32168 stainless steel were selected as welding materials, and manual arc welding and gas shielded welding processes were used to simulate actual production scenarios to prepare welding specimens; S32, artificially introduces defects such as cracks, pores, spatter and dents into the weld specimen to simulate the defects in the actual welding process; S33, V-shaped or U-shaped grooves are processed by CNC machine tools, and combined with laser cleaning to ensure the surface cleanliness of the welding specimens; the assembly accuracy of the welding specimens is guaranteed by fixture positioning and laser calibration, forming a structural form covering butt welds and fillet welds. The thickness of the welding specimens ranges from 12-20mm.

6. The method according to claim 1, wherein The secondary fine-tuning in step S4 includes the following steps: S41, by referencing the multi-layer perceptron module in the GLIP model, the appearance characteristics of the weld image and the time domain signal characteristics of the ultrasonic signal are cross-modally mapped; S42, unfreeze all frozen layers of the GLIP model, adjust the multilayer perceptron module parameters, including the number of channels and activation function, and adjust the training strategy, including the learning rate decay step size and gradient clipping threshold; S43: Build a fine-tuning configuration file dedicated to the target detection task, update the dataset path, switch the task type to target detection, and use Full Fine-tuning as the fine-tuning strategy. S44. Based on steps S41, S42 and S43, the GLIP model is fine-tuned for the second time in combination with a multimodal joint loss function. The joint loss function includes classification loss, regression loss and cross-modal consistency loss. The classification loss uses Focal Loss to alleviate the category imbalance problem, the regression loss uses CIOU Loss to accurately evaluate the positioning error of the target box, and the cross-modal consistency loss uses contrastive learning to constrain the multimodal features to be aligned in the shared semantic space.

Citation Information

Cited By

  • Welding structure defect labeling method and defect recognition model training method

    CN121391853A

  • Weld defect color image detection method and system based on YOLO-World

    CN121810699A