Multi-mode heterogeneous information collaborative weld defect X-ray image intelligent diagnosis and credible traceability method

By employing a multimodal collaborative intelligent diagnostic method for weld defects using X-ray images, and utilizing an intelligent diagnostic network for weld defects that integrates graph structure and multi-domain features, the method solves the problem that single image information is insufficient to comprehensively characterize defect features, and achieves stable identification and reliable traceability of weld defects.

CN121904024APending Publication Date: 2026-04-21TAIYUAN UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 0 Cites 2 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TAIYUAN UNIVERSITY OF SCIENCE AND TECHNOLOGY
Filing Date
2026-01-23
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing intelligent diagnostic methods for weld defects using X-ray images mostly rely on single image information, making it difficult to fully characterize the essential features of the defects. This results in limited stability and generalization ability of diagnostic results, and makes it difficult to trace the source of the detection results.

Method used

A multimodal heterogeneous information collaboration method is adopted, which takes weld X-ray images and industry inspection standard text as inputs to construct an intelligent weld defect diagnosis network based on graph structure and multi-domain feature fusion. The network propagates and aggregates weld image features through graph convolutional network and performs adaptive weighted fusion with text modal features to achieve weld defect type identification, severity assessment and compliance determination.

Benefits of technology

It enhances the stability and reliability of the weld defect diagnosis process, enables credible traceability of test results, and meets the intelligent and standardized requirements of industrial non-destructive testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121904024A_ABST
    Figure CN121904024A_ABST
Patent Text Reader

Abstract

The invention discloses a welding seam defect X-ray image intelligent diagnosis and credible traceability method based on multi-modal heterogeneous information collaboration, which is characterized in that a defect analysis network fusing multi-domain feature modeling and graph structure expression is constructed on the basis of bimodal data formed by a welding seam X-ray image and an industry detection standard text. In the image mode, dividing the weld seam image into a plurality of local area units through superpixel segmentation, taking the areas as image nodes, respectively extracting spatial domain, frequency domain, wavelet domain and edge domain features, and constructing a weighted graph structure by combining the spatial adjacency relation and the feature similarity relation between the areas; realizing overall modeling and correlation analysis of weld defect structure information by using a graph convolutional network; in a text mode, feature coding is carried out on an industry detection standard text, and the feature coding is used as an important prior constraint for defect judgment. Collaborative modeling of an image detection result and standard semantic information is achieved through a gating fusion mechanism, a mapping relation between a detection conclusion and a standard term is established, and interpretable expression and result credible traceability of the weld defect diagnosis process are achieved. And a welding seam X-ray film automatic digital acquisition and observation device is adopted in a matched manner, so that stable transmission, positioning observation and high-resolution digital imaging of the industrial ray film are realized, and reliable and consistent image data input is provided for the intelligent diagnosis method. The method is suitable for intelligent defect detection under complex welding seam structures and multi-working-condition imaging conditions, and has high engineering application value and popularization prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent non-destructive testing and industrial visual analysis technology, specifically a method for intelligent diagnosis and reliable tracing of weld defects using X-ray images based on multimodal heterogeneous information collaboration. Background Technology

[0002] In recent years, with the continuous expansion of my country's high-end equipment manufacturing, energy engineering, and major infrastructure construction, welded structures have been increasingly widely used in pressure vessels, pipeline systems, and key load-bearing components. Their welding quality directly affects the safety and reliability of engineering operations. Intelligent X-ray imaging diagnosis of weld defects, as a crucial supporting technology for the digital transformation of industrial inspection, faces challenges such as complex defect types, strong image noise interference, high reliance on manual interpretation, and difficulty in tracing the source of detection results. How to integrate multi-source heterogeneous information and construct a multi-modal collaborative intelligent diagnosis and reliable traceability method has become a key technical issue for improving the accuracy, reliability, and engineering application value of weld defect detection. In the field of intelligent diagnosis of weld defects using X-ray images, accurate identification and localization of defect types are key technologies for achieving objective assessment of welding quality and ensuring safety. The accuracy of the diagnostic results directly affects the service reliability and risk assessment conclusions of engineering structures. With the rapid development of computer vision and deep learning technologies, using target detection and image segmentation methods to automatically identify and discriminate defects in weld X-ray images has become a practical and feasible solution to replace traditional manual interpretation and improve detection efficiency and consistency. Wang Shusen et al. explored deep learning-based strategies for enhancing and recognizing weld defects in X-ray images, analyzing the impact of different image enhancement methods on recognition accuracy, and providing a theoretical and experimental basis for improving image quality preprocessing. He Yunkai et al. proposed a weld defect recognition method based on convolutional neural networks, combining data augmentation and transfer learning of pre-trained models on the GDX-ray dataset, achieving high classification performance. Mery et al. used feature reconstruction and cross-scale fusion methods on large-scale weld X-ray images to achieve a high level of automatic localization and improved detection accuracy. Liu et al. proposed a lightweight YOLO variant (LF-YOLO) for weld defect detection, which improves detection performance while ensuring real-time performance by enhancing multi-scale feature extraction and efficient feature encoding. The above research shows that deep learning-based intelligent diagnostic methods for weld defects using X-ray images have made some progress in defect identification and localization. However, existing studies mostly focus on single image information as the main analysis object, emphasizing the extraction of visual features at the pixel level, and failing to fully utilize prior information such as industry standards, inspection specifications, and working conditions involved in weld inspection. In complex industrial inspection scenarios, weld defects often exhibit diverse shapes, significant scale differences, and inconsistent imaging conditions. Relying solely on single-modal visual features is insufficient to comprehensively characterize the essential features of the defects, easily leading to limited stability and generalization ability of diagnostic results. Furthermore, existing methods have limited interpretability of the inspection basis during defect judgment, making it difficult to support the engineering reliability assessment of inspection results and subsequent quality management needs. To effectively address the aforementioned issues, this invention proposes a multimodal heterogeneous information collaborative intelligent diagnosis and reliable traceability method for weld defect X-ray images. This method, through multimodal collaborative training, jointly models weld X-ray image information with textual information such as industry standards and working condition descriptions. This improves the accuracy of intelligent defect diagnosis while establishing a correlation between the detection results and their formation conditions and judgment criteria. By collaboratively utilizing multimodal heterogeneous information, this method enhances the stability and reliability of the weld defect diagnosis process and provides traceable and interpretable support for the detection results, thereby better meeting the comprehensive application requirements of intelligent, standardized, and reliable industrial nondestructive testing. Summary of the Invention

[0003] To address the shortcomings of existing technologies, the present invention aims to provide a method for intelligent diagnosis and reliable tracing of weld defects using X-ray images based on multimodal heterogeneous information collaboration. This invention is achieved using the following technical solution: A method for intelligent diagnosis and reliable tracing of weld defects using X-ray images with multimodal heterogeneous information collaboration is proposed. The method uses bimodal data consisting of weld X-ray images and industry inspection standard text as input to construct an intelligent diagnostic network for weld defects based on graph structure and multi-domain feature fusion. In the image modality, superpixel segmentation is used to divide the weld X-ray image into regions, and each obtained locally consistent region is used as a node in the graph structure. For each node region, multi-domain visual features are extracted from the spatial domain, frequency domain, wavelet domain, and edge domain, and the multi-domain features are concatenated to form a node attribute representation. At the same time, spatial adjacency edges are constructed based on the physical spatial adjacency relationship between regions, and feature similarity edges are constructed based on the similarity relationship between the multi-domain features of nodes. By weighted fusion of spatial adjacency relationship and feature similarity relationship, a weighted graph structure reflecting the local structural features and global semantic consistency of the weld image is formed. Based on the graph structure, a graph convolutional network is used to propagate and aggregate node features, and graph pooling is used to obtain the image modality feature representation characterizing the overall structural features of weld defects. In the text modality, the industry testing standard text and the corresponding welding condition information are uniformly encoded using feature vectorization to obtain a priori semantic feature representation that simultaneously characterizes the weld defect judgment rules, compliance constraints, and imaging condition conditions. The condition information includes at least one of the welding method, material type, process parameters, and testing conditions. By employing a gated fusion mechanism to adaptively weight and fuse image modal features and text modal features, a comprehensive feature representation of weld defects constrained by both industry standards and operating conditions is generated. Based on this comprehensive feature, weld defect type identification, severity assessment, and compliance determination are completed. Simultaneously, a correlation mapping relationship is established between weld defect diagnosis results and corresponding industry testing standard clauses and welding condition information, enabling interpretable expression of the intelligent weld defect diagnosis process and reliable traceability of test results. This invention presents a self-constructed weld radiographic film scanning defect dataset, the WRF-DD dataset. It utilizes industrial radiographic films collected during actual weld inspection processes in enterprises, digitized using specialized scanning equipment to generate high-resolution weld image data. The overall data exhibits significant characteristics such as a realistic industrial background, complex imaging conditions, and diverse defect morphologies. The data acquisition process strictly adheres to current non-destructive testing standards for welds, objectively reflecting the imaging characteristics of weld defects under actual engineering inspection environments. Regarding the annotation method, this dataset adopts pixel-level semantic segmentation annotation. Professionals with experience in industrial non-destructive testing perform pixel-by-pixel fine annotation of the weld defect area according to relevant testing standards, ensuring the accuracy and consistency of the defect area outline. Semantic segmentation annotation not only includes the spatial location of the defect, but also accurately describes the morphological structure of the defect, which is beneficial for the model to learn the true geometric features and imaging patterns of the weld defect. During data acquisition and image acquisition, this invention employs an automated digital acquisition and observation device for weld X-ray film, which stably transmits, positions, and images industrial X-ray films acquired on-site. The device includes a film feeding mechanism for sequential film transport and positioning, an observation mechanism for film dwelling and uniform transmission illumination, and a digital reading unit for high-resolution imaging. It can automatically scan and digitize weld X-ray films while ensuring film flatness and imaging consistency. The weld X-ray images acquired by this device serve as the image modal input data for this invention, effectively ensuring consistency in spatial resolution, grayscale levels, and imaging stability, providing a reliable data foundation for subsequent intelligent diagnosis of weld defects based on graph structure and multi-domain feature fusion. This invention relates to an automatic digital acquisition and observation device for weld X-ray films. The device includes a viewing area, with a film delivery area on one side and a film receiving area on the other. A support frame is positioned above the viewing surface of the viewing area. Within the support frame, pulleys I, II, III, and IV, driven by a horizontal belt, are sequentially installed. A motor is mounted on the support frame near the film delivery area, driving a first longitudinal shaft located near the delivery area. Pulley I is mounted on the upper end of the first longitudinal shaft. An up-film transfer rubber roller is mounted on the upper part of the first longitudinal shaft (below pulley I), and a down-film transfer rubber roller is mounted on the lower part of the first longitudinal shaft. The up-film transfer rubber roller and the down-film transfer rubber roller are used to press the upper and lower edges of the weld X-ray film. Pulley II and the upper active guide roller I are coaxially mounted on the upper and lower ends of a second longitudinal shaft. A lower driven guide roller I is positioned below the upper active guide roller I. Pulley III... The upper active guide roller II is coaxially mounted on the upper and lower ends of the third longitudinal axis, and a lower driven guide roller II is provided below the upper active guide roller II; the pulley IV, the upper receiving rubber roller, and the lower receiving rubber roller are coaxially mounted on the fourth longitudinal axis; the upper sending rubber roller, the upper active guide roller I, the upper active guide roller II, and the upper receiving rubber roller are located on the same horizontal straight line, and the lower sending rubber roller, the lower driven guide roller I, the lower driven guide roller II, and the lower receiving rubber roller are located on the same horizontal straight line; two sensors are diagonally arranged in the observation area of ​​the film viewing area; the bottom surface of the observation area of ​​the film viewing area and the bottom surface of the receiving area are provided with a connected film guide groove, and the film guide groove is also connected to the front baffle of the film feeding area; a front baffle is fixedly provided at the front end of the bottom plate of the film feeding area, and a rear baffle is fixedly provided at the rear end of the bottom plate; a movable baffle is provided between the front baffle and the rear baffle, and the movable baffle is connected to the rear baffle by a compression spring; In use, the automatic digital acquisition and intelligent detection and evaluation method for weld X-ray films is as follows: The sample to be observed is placed between the front baffle and the movable baffle in the film delivery area. A compression spring keeps the sample pressed against the upper and lower transfer rubber rollers and the front baffle at all times. The motor is turned on, and the rotating upper and lower transfer rubber rollers separate the outermost sample from the other samples. Simultaneously, the belt transmits power synchronously to the upper active guide roller I, upper active guide roller II, and the upper and lower receiving rubber rollers, conveying the sample to the observation area of ​​the viewing area. When two sensors diagonally opposite each other in the observation area are simultaneously blocked by the sample, a signal is sent to the PLC control unit of the viewing lamp. The PLC control unit sends a signal to power off the motor and stop it, increasing the light intensity in the observation area and keeping the sample in the observation area for a certain period of time to perform the image acquisition process. After the reading device (viewer) in the observation area finishes reading, the PLC control unit ends the countdown and sends a power-on command to the motor. The upper and lower transfer rubber rollers continue to rotate and start transferring the next sample. At the same time, the sample remaining in the observation area is transferred to the upper and lower receiving rubber rollers by the active guide rollers and driven guide rollers located on both sides. Under the action of the film guide slot, the observed sample falls into the receiving area, completing the entire automatic observation process of the sample. This invention is rationally designed and applicable to X-ray non-destructive testing scenarios with complex weld structures and multiple working conditions. It can intelligently identify and comprehensively analyze the type, spatial distribution, and severity of weld defects, and realize the correlation mapping between test conclusions and industry standard clauses. It provides a reliable basis for welding quality assessment, engineering safety evaluation, and traceability of test results, effectively improving the intelligence, standardization, and reliability of the weld inspection process, and helping to upgrade the digitalization and intelligence of industrial non-destructive testing. It has good engineering application prospects and promotion value. Detailed Implementation

[0004] Many specific details are set forth in the following description in order to provide a full understanding of the invention, but the invention may also be practiced in other ways different from those described herein; obviously, the embodiments in the specification are only some embodiments of the invention, and not all embodiments. The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings; A method for intelligent diagnosis and reliable source tracing of weld defects using X-ray images based on multimodal heterogeneous information collaboration, such as... Figure 2 As shown, a dual-modal data consisting of weld X-ray images and industry inspection standard texts is used as input to construct an intelligent weld defect diagnosis network based on graph structure and multi-domain feature fusion. In the image modality, superpixel segmentation is used to divide the weld X-ray image into regions, and each obtained locally consistent region is used as a node in the graph structure. For each node region, multi-domain visual features are extracted from the spatial domain, frequency domain, wavelet domain, and edge domain, and the multi-domain features are concatenated to form a node attribute representation. At the same time, spatial adjacency edges are constructed based on the physical spatial adjacency relationship between regions, and feature similarity edges are constructed based on the similarity relationship between the multi-domain features of nodes. By weighted fusion of spatial adjacency relationship and feature similarity relationship, a weighted graph structure reflecting the local structural features and global semantic consistency of the weld image is formed. Based on the graph structure, a graph convolutional network is used to propagate and aggregate node features, and graph pooling is used to obtain the image modality feature representation characterizing the overall structural features of weld defects. In the text modality, the industry testing standard text and corresponding welding condition information are uniformly encoded using feature vectorization, such as... Figure 3 As shown, a priori semantic feature representation that simultaneously characterizes weld defect judgment rules, compliance constraints, and imaging conditions is obtained, wherein the condition information includes at least one of welding method, material type, process parameters, and detection conditions. By employing a gated fusion mechanism to adaptively weight and fuse image modal features and text modal features, a comprehensive feature representation of weld defects constrained by both industry standards and operating conditions is generated. Based on this comprehensive feature, weld defect type identification, severity assessment, and compliance determination are completed. Simultaneously, a correlation mapping relationship is established between weld defect diagnosis results and corresponding industry testing standard clauses and welding condition information, enabling interpretable expression of the intelligent weld defect diagnosis process and reliable traceability of test results. The implementation details of the graph-based multi-domain feature modeling method are as follows: The input weld X-ray image is first subjected to superpixel segmentation, dividing the original image into several locally consistent regional units, and each region is used as a node in the graph structure. For each node region, the minimum bounding rectangle image block of the corresponding region is extracted and scale normalized. Multi-domain visual features are extracted from the spatial domain, frequency domain, wavelet domain, and edge domain, respectively. The spatial domain features are used to describe the gray-level distribution and texture statistical characteristics, the frequency domain features are used to characterize the imaging spectrum distribution characteristics, the wavelet domain features are used to depict multi-scale structural information, and the edge domain features are used to enhance the expression of defect contour and orientation features. The features of each domain are stitched together to form the multi-domain attribute feature vector of the node, thereby realizing a comprehensive description of the local structure and imaging characteristics of the weld defect. Furthermore, spatial adjacency edges between regions are constructed based on the physical spatial contact relationship of each superpixel region in the weld X-ray image. Simultaneously, cross-regional feature similarity edges are constructed based on the similarity relationship between multi-domain features of nodes. The spatial adjacency edges and feature similarity edges are then weighted and fused to form a weighted adjacency matrix. The node multi-domain attribute feature vectors and the weighted adjacency matrix are input into a graph convolutional network. Multi-layer graph convolution operations are used to realize the propagation and aggregation of node features on the graph structure. During feature propagation, normalization processing and residual connection mechanisms are used to enhance the model training stability. Finally, node-level features are weighted and aggregated through graph pooling based on an attention mechanism to obtain a global feature representation of the image modality representing the overall structural information of weld defects, which is used for subsequent multi-modal collaborative fusion and intelligent diagnosis of weld defects. The aforementioned image modality is modeled using multi-domain pixel-aware superpixel nodes, and its model structure diagram is as follows: Figure 4 As shown, the specific construction process is as follows: The image modality adopts a multi-domain pixel-aware superpixel node modeling method. Its core design idea is to replace pixel-by-pixel modeling with regional-level feature abstraction, so that the structural information of weld defects can be more stably expressed in X-ray images with large noise interference. In the specific implementation process, the input weld X-ray image is first subjected to superpixel segmentation, dividing the original image into several regions with consistent grayscale distribution, texture statistics, and spatial continuity. Each superpixel region is used as a node in the graph structure to avoid the high sensitivity of pixel-level modeling to local noise, uneven exposure, and scattering artifacts. For the i-th superpixel node, a graph structure is constructed as follows: Figure 5 The multi-domain feature extractor shown below has its multi-domain attribute feature vectors expressed as follows: ; Among them, spatial features The local imaging characteristics of the weld area are characterized by grayscale statistics and texture statistics, which are used to distinguish the grayscale stability differences between the defect area and the background area; frequency domain features Spectral statistical features were extracted using Fourier transform to reflect imaging frequency perturbations caused by defects; wavelet domain features were also used. Multi-scale wavelet decomposition is used to characterize the structural changes of defects at different spatial scales; marginal domain features. Enhance defect contour, directionality, and boundary discontinuity information through edge detection and gradient direction analysis; The fundamental reason for adopting multi-domain feature joint modeling is: Weld defects in X-ray images are usually not manifested as a single feature anomaly, but rather as the result of the superposition of multiple imaging characteristics; by stitching together multi-domain features, each node can have a more comprehensive source of discriminative information during the training process; During the model training phase, multi-domain features are used as the initial input to the graph convolutional network. During backpropagation, they are jointly constrained by loss functions such as defect classification, severity assessment, and compliance prediction, enabling the model to adaptively learn the relative importance of each feature domain for weld defect identification, thereby improving the overall detection accuracy and robustness. The spatial adjacency graph based on pixel-level physical contact relationships is constructed as follows: By explicitly modeling the local continuity and structural correlation of weld defects in X-ray images, the physical imaging prior of weld defects is constructed to construct the structural propagation process, fundamentally constraining the rationality of feature propagation. In weld X-ray images, defects typically exhibit a continuous distribution along the weld direction, such as linear crack propagation, banded distribution of incomplete fusion defects, and localized aggregation of porosity. These defects demonstrate significant continuity in physical space, and relying solely on feature similarity is susceptible to cross-regional semantic interference that does not conform to the physical structure. Therefore, this method first establishes spatial adjacency constraints through pixel-level spatial relationships. Specifically, the spatial adjacency matrix is ​​defined as: ; Among them, pixel contact relationships are determined by scanning four or eight neighboring areas; The purpose of this formula is to allow only physically adjacent regions to establish basic connections in the graph structure, thereby ensuring the consistency between the graph structure topology and the weld imaging structure. During the graph convolution propagation stage, the spatial adjacency matrix serves as a basic constraint, ensuring that node features primarily propagate within their local neighborhoods. This strengthens the structural integrity of weld defects and prevents features from crossing non-adjacent regions during propagation, thus avoiding structural damage. During model training, this spatial constraint mechanism enables the network to prioritize learning discrimination patterns consistent with the weld geometry, significantly reducing the probability of false detections caused by noise, uneven exposure, or imaging artifacts, and improving the stability and reliability of the model in complex weld scenarios. The cross-regional connection mechanism for multi-domain feature similarity is as follows: To address the issue that relying solely on spatial adjacency relationships is insufficient to characterize the semantic consistency of weld defects across the entire image, a cross-regional connection mechanism based on multi-domain feature similarity is designed to enhance the overall perception capability of scattered, repetitive, and weak-contrast defects. In weld X-ray images, some defects, although spatially discontinuous, exhibit high similarity in texture, spectrum, or structural features, such as porosity, slag inclusions, or microcracks distributed across different weld locations. Relying solely on spatial adjacency relationships is insufficient for effective information exchange between these defect regions; therefore, this method is based on node multi-domain feature vectors. The construction of feature similarity relationships is defined as follows: ; The similarity function uses cosine similarity: ; Multi-domain features have differences in dimensions and amplitudes across different dimensions. Cosine similarity can focus on the consistency of feature distribution patterns rather than simply numerical magnitude, making it more suitable for measuring the similarity of weld defects at the imaging pattern level. By constructing connections between several nodes that are most similar to each node in terms of its features, regions with similar defect features can establish long-distance semantic connections in the graph structure, thereby achieving cross-region feature enhancement during graph convolution propagation. During the training phase, this mechanism enables the model to learn the global discrimination pattern of weld defects, improving the overall detection capability for small defects, low-contrast defects, and discretely distributed defects. The specific details of the weighted fusion of spatial adjacency and feature similarity are as follows: By weighted fusion of spatial adjacency relationships and feature similarity relationships, a unified weighted adjacency matrix is ​​constructed to achieve the collaborative propagation of local structural information and global semantic information of weld defects; the weighted adjacency matrix is ​​defined as follows: ; in, The physical rationale used to constrain feature propagation; Used to enhance cross-regional semantic consistency; This fusion approach addresses the problem that single spatial adjacency can lead to an overemphasis on local structures and neglect of global defect patterns, while also addressing the issue that similar adjacencies of a single feature may disrupt the spatial topological consistency of weld imaging. Through weighting coefficients and This enables the model to achieve a dynamic balance between "structural constraints" and "semantic enhancement" during training. During the training phase, the weighted adjacency matrix participates in the end-to-end backpropagation of the graph convolutional network. The weight parameters can be used as fixed hyperparameters or tuned through the validation set, thereby enabling the model to maintain good generalization performance under different weld types and imaging conditions. The details of the residual connections and normalized multi-layer graph convolutional stable propagation and deep feature learning mechanisms used are as follows: The graph convolutional network employs a multi-layer stacked structure and performs residual connections and normalization operations during feature propagation to address the common problems of over-smoothing features and training instability in deep graph convolutional networks. The feature update form of the t-th layer graph convolution is as follows: ; in, Let A represent the node feature matrix of the t-th layer; let A represent the adjacency matrix of the graph. LN represents the learnable weight matrix of the t-th layer; LN denotes the layer normalization operation. This represents the activation function, such as ReLU, GELU, etc. residuals This is used to preserve the original node features and prevent node features from becoming homogenized during multi-layer propagation; the layer normalization operation is used to constrain the feature distribution and reduce the interference of imaging condition differences between different weld samples on the training process; this structural design enables the model to safely stack multi-layer graph convolutions, thereby capturing the local structural features, cross-regional semantic relationships and global discrimination patterns of weld defects layer by layer during training. The attention-based graph pooling approach is implemented as follows: By designing a graph pooling method based on an attention mechanism, node-level features are weighted and aggregated to generate a global structural feature representation of weld defects; the attention weight calculation method is defined as follows: ; in, Let represent the node-level feature vector of the i-th node after propagation in the graph convolutional network, and w be a learnable linear mapping parameter matrix used to map the node features. Mapped to the attention feature space; the final graph-level feature representation is: ; The purpose of this attention mechanism is to solve the problem of "sparse distribution of key information" in weld images. By learning node weights, regions containing obvious defect features dominate the global features, while interference from background regions is effectively suppressed. During training, the attention weights are automatically adjusted with backpropagation of the loss function, enabling the model to gradually develop the ability to explicitly focus on key defect regions, thereby improving the accuracy and interpretability of the overall diagnostic results.

[0005] The gated fusion mechanism used for adaptive fusion of image modal features and text modal features is as follows: Figure 6 As shown, the specific implementation details are as follows: Image modal features and text modal features are adaptively fused using a gated fusion mechanism and jointly trained under a unified optimization objective, thereby achieving reliable traceability of weld defect diagnosis results. The gated fusion method is defined as follows: ; in, This represents the text modal features after encoding industry testing standards and welding condition information; This represents the structural features of weld defects extracted from image modal analysis; The mechanism of this design is as follows: when image evidence is insufficient or defect features are not obvious, the model can automatically enhance the influence of industry standard semantics; when image features are clear and reliable, the model increases the image modal weights, making the diagnostic results more dependent on visual evidence. During the training phase, the gating weights are jointly optimized along with the task loss function, and through methods such as... Figure 7 The defect classification detection head shown performs multi-scale joint discrimination and regression on the fused features, enabling the weld defect type determination, severity assessment and compliance analysis results to establish a correspondence with specific industry standard clauses, thereby achieving interpretable output of the diagnostic process and reliable traceability of results; The intelligent diagnosis and reliable source tracing method for weld defects using X-ray images based on multimodal heterogeneous information collaboration uses a self-constructed WRF-DD dataset for model pre-training. By jointly learning weld X-ray images and their corresponding defect annotation information, the network can fully grasp the typical imaging features, structural distribution patterns and category discrimination information of weld defects in the initial stage, thus providing a stable and reliable feature representation foundation for subsequent multimodal collaborative diagnosis and reliable source tracing tasks. The self-constructed WRF-DD dataset uses industrial radiographic films collected during actual weld inspections in enterprises. These films are digitized using professional scanning equipment to form high-resolution weld image data. The overall data is characterized by a realistic industrial background, complex imaging conditions, and diverse defect morphologies. The data acquisition process strictly follows current non-destructive testing standards for welds, and can objectively reflect the imaging characteristics of weld defects under actual engineering inspection environments. Regarding the annotation method, this dataset adopts pixel-level semantic segmentation annotation. Professionals with experience in industrial non-destructive testing perform pixel-by-pixel fine annotation of the weld defect area according to relevant testing standards, ensuring the accuracy and consistency of the defect area outline. Semantic segmentation annotation not only includes the spatial location of the defect, but also accurately describes the morphological structure of the defect, which is beneficial for the model to learn the true geometric features and imaging patterns of the weld defect. Automatic digital acquisition and observation equipment for X-ray film of welds, such as Figure 8 As shown, it includes viewing area 1 (including film viewer and film transfer mechanism), film delivery area 2 (including film delivery mechanism) and receiving area 3 (including receiving rack). Area 101 is a light shield, and its internal structure is as follows: Figure 9 As shown, viewing area 1 has a film delivery area 2 on one side and a film receiving area 3 on the other side; it is used to automatically realize the film delivery, film acquisition after the film is in place, and film receiving while ensuring the original numbering order of the weld X-ray film. like Figure 9 , 10As shown, a light shield 101 is provided above the observation surface of the observation area 1. Inside the light shield 101, pulleys I 107, II 108, III 109, and IV 110, driven by belts 111, are horizontally installed in sequence. (Pulleys I 107 and IV 110 are symmetrically arranged about the center line of the observation surface, and pulleys II 108 and III 109 are also symmetrically arranged about the center line of the observation surface.) A motor 102 is installed on the light shield 101 near the film delivery area 2. The motor 102 drives a first longitudinal shaft 103 located near the film delivery area 2. Pulley I 107 is mounted on the first longitudinal shaft 103. At the upper end, a first longitudinal shaft 103 has an upper feed rubber roller 112 mounted on its upper part and a lower feed rubber roller 113 mounted on its lower part. The upper feed rubber roller 112 and the lower feed rubber roller 113 are used to press the upper and lower edges of the weld X-ray film. A pulley II 108 and an upper active guide roller I 114 are coaxially mounted on the upper and lower ends of the second longitudinal shaft 104. A lower driven guide roller I 115 is located below the upper active guide roller I 114. A pulley III 109 and an upper active guide roller II 116 are coaxially mounted on the upper and lower ends of the third longitudinal shaft 105. The upper active guide roller II 116 is located below... The upper driven guide roller II 117 is provided; the pulley IV 110, the upper receiving rubber roller 118, and the lower receiving rubber roller 119 are coaxially mounted on the fourth longitudinal axis 106; the upper receiving rubber roller 112, the upper active guide roller I 114, the upper active guide roller II 116, and the upper receiving rubber roller 118 are located on the same horizontal straight line (the upper receiving rubber roller 112 and the upper receiving rubber roller 118 are symmetrically arranged about the center line of the observation surface, and the upper active guide roller I 114 and the upper active guide roller II 116 are symmetrically arranged about the center line of the observation surface), and the lower receiving rubber roller 113... The lower driven guide roller I 115, lower driven guide roller II 117, and lower receiving rubber roller 119 are located on the same horizontal straight line (the lower transfer rubber roller 113 and the lower receiving rubber roller 119 are symmetrically arranged about the center line of the observation surface, and the lower driven guide roller I 115 and the lower driven guide roller II 117 are symmetrically arranged about the center line of the observation surface); two sensors 120 are diagonally arranged in the observation area of ​​observation area 1; the bottom surface of the observation area of ​​observation area 1 and the bottom surface of the receiving area 3 are provided with a connected film guide groove 121, which is also connected to the front baffle 201 of the film delivery area; Figure 11 , 12 As shown, a front baffle 201 is fixedly installed at the front end of the base plate of the feeding area 2, and a rear baffle 202 is fixedly installed at the rear end of the base plate. A movable baffle 203 is located between the front baffle 201 and the rear baffle 202, and a compression spring 204 is connected between the movable baffle 203 and the rear baffle 202. The sample to be observed is placed between the front baffle 201 and the movable baffle 203 of the feeding and receiving mechanism. The movable baffle 203 is connected to the compression spring 204 to keep the sample pressed against the upper and lower transfer rubber rollers and the front baffle at all times. Figure 12As shown, the front baffle is provided with a film guide groove (L-shaped guide groove) 121 facing the observation area, and the lower edge of the light shield 101 is also provided with a corresponding film guide groove (L-shaped guide groove) to ensure that the sample does not fall off the viewer or rollers during the transmission process. A sample to be observed is placed between the front baffle 201 and the movable baffle 203 in the sample delivery area. A compression spring 204 keeps the sample pressed against the upper and lower transfer rubber rollers 112 and 113 and the front baffle. The motor 102 is turned on, transmitting power to the upper and lower transfer rubber rollers. The rotating rollers separate the outermost sample from the other samples. Simultaneously, the belt 111 transmits power to the upper active guide roller I 114, upper active guide roller II 116, and the upper and lower receiving rubber rollers 118 and 119, conveying the sample to the observation area of ​​observation area 1. When two sensors 120 diagonally opposite each other in the observation area are simultaneously blocked by the sample, a signal is sent to the PLC control unit. The PLC control unit then sends a signal to de-energize and stop the motor 102, increasing the light intensity in the observation area and keeping the sample in the observation area for a certain period of time to perform the image acquisition process. After the reading device finishes reading, the PLC control unit ends the countdown and sends a power-on command to the motor 102. The upper and lower transfer rubber rollers 112 and 113 continue to rotate to start transferring the next sample. At the same time, the sample remaining in the observation area is transferred to the upper and lower receiving rubber rollers 118 and 119 under the action of the active guide rollers and driven guide rollers located on both sides. Under the action of the film guide slot 121, the observed sample falls into the receiving area, completing the entire automatic observation process of the sample. The image acquisition process is as follows: The film reading device of the film viewing machine in the viewing area acquires digital images of the weld X-ray film that has been transmitted and placed in position through an 8K ultra-high-definition camera 205; the 8K ultra-high-definition camera 205 should be positioned and set directly towards the viewing area to acquire digital images of the weld X-ray film that has been transmitted and placed in position, and is connected to the PC host computer; during operation, when the film is in position, the PLC control of the automatic film viewing light device stops the film transmission, enhances the light intensity in the viewing area, and transmits feedback information to the PC in real time. Based on the feedback information from the PLC, the PC promptly calls the UVC protocol through OpenCV to start controlling the 8K ultra-high-definition camera 205 to acquire digital images of the film frame by frame. Each film is acquired at intervals set through the interactive operation interface GUI, frame by frame. When two acquired digital images of the film are the same, the acquisition result of the previous X-ray film is overwritten. After the digital film information acquisition and analysis are completed, the PC controls the camera to stop recording and sends a start transmission command to the PLC of the automatic film viewing light device to continue the transmission and update of the film, completing the camera's acquisition process of the film. It also supports manual control of data acquisition; The PC-based host computer is mainly used for controlling the film viewing light device and camera, coordinating the recognition and calculation of various components, data storage, and the graphical user interface. This invention was conducted under the conditions of an Intel(R) Core(TM) i5-13400F CPU, an NVIDIA GeForce RTX 4070 TiSUPER (32GB), a 64-bit Windows operating system, and PyCharm 2022.2.1. The pre-training used image modality multi-domain feature dimensions are shown in Table 1. The overall configuration and parameter description of the defect detection bimodal model are shown in Table 2. The experimental hyperparameters and system configuration description are shown in Table 3. Table 1. Composition of Multi-Domain Feature Dimensions of Image Modalities Feature Domain Type Feature Dimension airspace features 17-dimensional Frequency domain characteristics 9D Wavelet domain features 21-dimensional marginal domain features 11-dimensional Total Dimension of Multi-Domain Features 58-dimensional Table 2 Overall Configuration and Parameter Description of the Dual-Modal Defect Detection Model project Content Description Model Name Intelligent Weld Defect Diagnosis Network Based on Graph Structure and Multi-Domain Feature Fusion Modal composition Image modality (graph structure + multi-domain features) + text modality (industry detection standards + operating condition information) Model parameter count 579, 601 Image feature types Spatial domain features + Frequency domain features + Wavelet domain features + Edge domain features Detect output content Defect type (5 categories) + severity (4 levels) + compliance assessment + defect boundary box Table 3 Experimental Hyperparameters and System Configuration Description category parameter numerical values Input settings Input mode Text, images Image size 494×494 Model Structure Hidden feature dimensions 256 Text embedding dimension 768 Structural parameters Mamba layers 4 GCN layers 2 Training parameters BatchSize 128 Number of training rounds 300 Learning rate <![CDATA[1 × 10 -4 ]]> Weight decay <![CDATA[1 × 10 -5 ]]> In the actual inspection process, the system uses weld X-ray images as the sole input data. These images are stably transmitted, positioned, observed, and digitally acquired at high resolution by an automated digital acquisition and observation device for weld X-ray film. Based on the above image input, the model can automatically output weld defect detection results and generate a structured inspection report. The report covers multiple results, including defect type determination, severity assessment, defect spatial location, geometric dimension analysis, and compliance assessment. The specific output results are shown below. Weld Defect Inspection Report I. Report Information Report Number: 20260107151754 Detection time: 2026-01-07 15:17:54 Testing standard: GB / T 3323-2005 Detection method: Intelligent vision inspection system II. Sample Information Sample number: 1-12 Welding method: Manual arc welding Base material: Q345 steel plate Sheet thickness: 12mm Image file: 1-12.png Image size: 793 × 662 pixels III. Test Results 3.1 Defect Classification Defect type: Crack Recognition confidence level: 62.48% Severity: Moderate Confidence level: 71.32% 3.2 Defect Location Center coordinates: (283, 381) Bounding box: Top left (59, 162) - Bottom right (507, 599) Relative location: Central 3.3 Defect Dimensions Width: 448 pixels Height: 437 pixels Area: 195,776 square pixels Aspect ratio: 1.03 IV. Probability Analysis 4.1 Defect Type Probability Distribution Cracks: 62.48% Incomplete penetration: 14.27% Slag inclusions: 10.36% Edge bite: 6.15% Stomata: 6.74% 4.2 Severity Probability Distribution Mild: 12.84% Medium: 71.32% Severe: 8.46% Extremely severe: 7.38% V. Compliance Assessment Compliance score: 0.5111 Evaluation threshold: 0.7 Assessment result: Unsatisfactory VI. Test Results [Conclusion] Cracks were detected, with a moderate severity. [Recommendation] The defect is of moderate severity. It is recommended to strengthen monitoring and assess whether rework is necessary. VII. Additional Notes 1. This report was automatically generated based on an intelligent visual inspection system. 2. Test results are for reference only; final judgment requires manual review. 3. Confidence level indicates the reliability of the model's predictions, ranging from 0% to 100%. 4. The defect location coordinate system has its origin at the top left corner of the image. Report generated: January 7, 2026, 15:17:54 [End of report] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the present invention. Although detailed descriptions have been made with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments, and they should all be covered within the protection scope of the claims.

Claims

1. A method for intelligent diagnosis and reliable tracing of weld defects using X-ray images based on multimodal heterogeneous information collaboration, characterized in that: Using dual-modal data consisting of weld X-ray images and industry inspection standard texts as input, an intelligent weld defect diagnosis network based on graph structure and multi-domain feature fusion is constructed. In the image modality, superpixel segmentation is used to divide the weld X-ray image into regions, and each obtained locally consistent region is used as a node in the graph structure. For each node region, multi-domain visual features are extracted from the spatial domain, frequency domain, wavelet domain, and edge domain, and the multi-domain features are concatenated to form a node attribute representation. At the same time, spatial adjacency edges are constructed based on the physical spatial adjacency relationship between regions, and feature similarity edges are constructed based on the similarity relationship between the multi-domain features of nodes. By weighted fusion of spatial adjacency relationship and feature similarity relationship, a weighted graph structure reflecting the local structural features and global semantic consistency of the weld image is formed. Based on the graph structure, a graph convolutional network is used to propagate and aggregate node features, and graph pooling is used to obtain the image modality feature representation characterizing the overall structural features of weld defects. In the text modality, the industry testing standard text and the corresponding welding condition information are uniformly encoded using feature vectorization to obtain a priori semantic feature representation that simultaneously characterizes the weld defect judgment rules, compliance constraints, and imaging condition conditions. The condition information includes at least one of the welding method, material type, process parameters, and testing conditions. By employing a gated fusion mechanism to adaptively weight and fuse image modal features and text modal features, a comprehensive feature representation of weld defects constrained by both industry standards and operating conditions is generated. Based on these comprehensive features, weld defect type identification, severity assessment, and compliance determination are completed. Simultaneously, a correlation mapping relationship is established between weld defect diagnosis results and corresponding industry testing standard clauses and welding condition information, enabling interpretable expression of the intelligent weld defect diagnosis process and reliable traceability of test results.

2. The intelligent diagnosis and reliable tracing method for weld defects using X-ray images based on multimodal heterogeneous information collaboration as described in claim 1, characterized in that: The image modality modeling employs a graph-based multi-domain feature map convolutional network, the specific construction process of which is as follows: A graph-based multi-domain feature modeling approach is adopted. The input weld X-ray image is first subjected to superpixel segmentation, dividing the original image into several locally consistent regional units, and each region is used as a node in the graph structure. For each node region, the minimum bounding rectangle image block of the corresponding region is extracted and scale normalized. Multi-domain visual features are extracted from the spatial domain, frequency domain, wavelet domain, and edge domain. The spatial domain features are used to describe the gray-level distribution and texture statistics, the frequency domain features are used to characterize the imaging spectrum distribution, the wavelet domain features are used to depict multi-scale structural information, and the edge domain features are used to enhance the expression of defect contour and orientation features. The features of each domain are concatenated to form the multi-domain attribute feature vector of the node, thereby realizing a comprehensive description of the local structure and imaging characteristics of the weld defect. Furthermore, spatial adjacency edges between regions are constructed based on the physical spatial contact relationships of each superpixel region in the weld X-ray image. Simultaneously, cross-regional feature similarity edges are constructed based on the similarity relationships between multi-domain features of nodes. The spatial adjacency edges and feature similarity edges are then weighted and fused to form a weighted adjacency matrix. The node multi-domain attribute feature vectors and the weighted adjacency matrix are input into a graph convolutional network. Multi-layer graph convolution operations are used to achieve the propagation and aggregation of node features on the graph structure. During feature propagation, normalization processing and residual connection mechanisms are used to enhance the model training stability. Finally, node-level features are weighted and aggregated using graph pooling based on an attention mechanism to obtain a global feature representation of the image modality representing the overall structural information of weld defects. This representation is used for subsequent multi-modal collaborative fusion and intelligent diagnosis of weld defects.

3. The intelligent diagnosis and reliable tracing method for weld defects using X-ray images based on multimodal heterogeneous information collaboration as described in claim 2, characterized in that: The image modality is modeled using a multi-domain pixel-aware superpixel node approach: The image modality adopts a multi-domain pixel-aware superpixel node modeling method. Its core design idea is to replace pixel-by-pixel modeling with regional-level feature abstraction, so that the structural information of weld defects can be more stably expressed in X-ray images with large noise interference. In the specific implementation process, the input weld X-ray image is first subjected to superpixel segmentation, dividing the original image into several regions with consistent grayscale distribution, texture statistics, and spatial continuity. Each superpixel region is used as a node in the graph structure to avoid the high sensitivity of pixel-level modeling to local noise, uneven exposure, and scattering artifacts. For the i-th superpixel node, the following multi-domain attribute feature vector is constructed: ; Among them, spatial features The local imaging characteristics of the weld area are characterized by grayscale statistics and texture statistics, which are used to distinguish the grayscale stability differences between the defect area and the background area; frequency domain features Spectral statistical features are extracted using Fourier transform to reflect imaging frequency perturbations caused by defects; wavelet domain features. Multi-scale wavelet decomposition is used to characterize the structural changes of defects at different spatial scales; marginal domain features. Enhance defect contour, directionality, and boundary discontinuity information through edge detection and gradient direction analysis; The fundamental reason for adopting multi-domain feature joint modeling is: Weld defects in X-ray images are usually not manifested as a single feature anomaly, but rather as the result of the superposition of multiple imaging characteristics; by stitching together multi-domain features, each node can have a more comprehensive source of discriminative information during the training process; During the model training phase, multi-domain features serve as the initial input to the graph convolutional network. During backpropagation, they are jointly constrained by loss functions such as defect classification, severity assessment, and compliance prediction, enabling the model to adaptively learn the relative importance of each feature domain for weld defect identification, thereby improving the overall detection accuracy and robustness.

4. The intelligent diagnosis and reliable tracing method for weld defects using X-ray images based on multimodal heterogeneous information collaboration as described in claim 2, characterized in that: The image modality is constructed by building a spatial adjacency graph based on pixel-level physical contact relationships: By explicitly modeling the local continuity and structural correlation of weld defects in X-ray images, the physical imaging prior of weld defects is constructed to construct the structural propagation process, fundamentally constraining the rationality of feature propagation. In weld X-ray images, defects typically exhibit a continuous distribution along the weld direction, such as linear crack propagation, banded distribution of non-fusion defects, and localized aggregation of pores. These defects have obvious continuity in physical space, and relying solely on feature similarity is easily subject to cross-regional semantic interference that does not conform to the physical structure. Therefore, this method first establishes spatial adjacency constraints through pixel-level spatial relationships. Specifically, the spatial adjacency matrix is ​​defined as: ; Among them, pixel contact relationships are determined by scanning four or eight neighboring areas; The purpose of this formula is to allow only physically adjacent regions to establish basic connections in the graph structure, thereby ensuring the consistency between the graph structure topology and the weld imaging structure. During the graph convolution propagation stage, the spatial adjacency matrix serves as a basic constraint, ensuring that node features primarily propagate within their local neighborhoods. This strengthens the structural integrity of weld defects and prevents features from crossing non-adjacent regions during propagation, thus avoiding structural damage. During model training, this spatial constraint mechanism enables the network to prioritize learning discrimination patterns consistent with the weld geometry, significantly reducing the probability of false detections caused by noise, uneven exposure, or imaging artifacts, and improving the stability and reliability of the model in complex weld scenarios.

5. The intelligent diagnosis and reliable tracing method for weld defects using X-ray images based on multimodal heterogeneous information collaboration as described in claim 2, characterized in that: Cross-region connection mechanism based on multi-domain feature similarity: To address the issue that relying solely on spatial adjacency relationships is insufficient to characterize the semantic consistency of weld defects across the entire image, a cross-regional connection mechanism based on multi-domain feature similarity is designed to enhance the overall perception capability of scattered, repetitive, and weak-contrast defects. In weld X-ray images, some defects, although spatially discontinuous, exhibit high similarity in texture, spectrum, or structural features, such as porosity, slag inclusions, or microcracks distributed across different weld locations. Relying solely on spatial adjacency is insufficient for effective information exchange between these defect regions. Therefore, this method is based on node multi-domain feature vectors. The construction of feature similarity relationships is defined as follows: ; The similarity function uses cosine similarity: ; Multi-domain features have differences in dimensions and amplitudes across different dimensions. Cosine similarity can focus on the consistency of feature distribution patterns rather than simply numerical magnitude, making it more suitable for measuring the similarity of weld defects at the imaging pattern level. By constructing connections between several nodes that are most similar to each node in terms of its features, regions with similar defect features can establish long-distance semantic connections in the graph structure, thereby achieving cross-region feature enhancement during graph convolution propagation. During the training phase, this mechanism enables the model to learn the global discrimination pattern of weld defects, improving the overall detection capability for small defects, low-contrast defects, and discretely distributed defects.

6. The intelligent diagnosis and reliable tracing method for weld defects using X-ray images based on multimodal heterogeneous information collaboration as described in claim 2, characterized in that: Weighted fusion of spatial adjacency and feature similarity: By weighted fusion of spatial adjacency relationships and feature similarity relationships, a unified weighted adjacency matrix is ​​constructed to achieve the collaborative propagation of local structural information and global semantic information of weld defects; the weighted adjacency matrix is ​​defined as follows: ; in, The physical rationale used to constrain feature propagation; Used to enhance cross-regional semantic consistency This fusion approach addresses the problem that single spatial adjacency can lead to an overemphasis on local structures and neglect of global defect patterns, while also addressing the issue that similar adjacencies of a single feature may disrupt the spatial topological consistency of weld imaging. Through weighting coefficients and This allows the model to achieve a dynamic balance between "structural constraints" and "semantic enhancement" during training. During the training phase, the weighted adjacency matrix participates in the end-to-end backpropagation of the graph convolutional network. The weight parameters can be used as fixed hyperparameters or tuned through the validation set, thereby enabling the model to maintain good generalization performance under different weld types and imaging conditions.

7. The intelligent diagnosis and reliable tracing method for weld defects using X-ray images based on multimodal heterogeneous information collaboration as described in claim 1, characterized in that: Residual connections and normalized multi-layer graph convolutional stable propagation and deep feature learning mechanism: The graph convolutional network employs a multi-layer stacked structure and performs residual connections and normalization operations during feature propagation to address the common problems of over-smoothing features and training instability in deep graph convolutional networks. The feature update form of the t-th layer graph convolution is as follows: ; in, Let A represent the node feature matrix of the t-th layer; let A represent the adjacency matrix of the graph. LN represents the learnable weight matrix of the t-th layer; LN denotes the layer normalization operation. This refers to activation functions, such as ReLU, GELU, etc. residuals This is used to preserve the original node features and prevent node features from becoming homogenized during multi-layer propagation; the layer normalization operation is used to constrain the feature distribution and reduce the interference of imaging condition differences between different weld samples on the training process; this structural design enables the model to safely stack multi-layer graph convolutions, thereby capturing the local structural features, cross-regional semantic relationships and global discrimination patterns of weld defects layer by layer during training.

8. The intelligent diagnosis and reliable tracing method for weld defects using X-ray images based on multimodal heterogeneous information collaboration as described in claim 1, characterized in that: Attention-based graph pooling: By designing a graph pooling method based on an attention mechanism, node-level features are weighted and aggregated to generate a global structural feature representation of weld defects; the attention weight calculation method is defined as follows: ; in, Let represent the node-level feature vector of the i-th node after propagation in the graph convolutional network, and w be a learnable linear mapping parameter matrix used to map the node features. Mapped to the attention feature space. The final graph-level feature representation is: ; The purpose of this attention mechanism is to solve the problem of "sparse distribution of key information" in weld images. By learning node weights, regions containing obvious defect features dominate the global features, while interference from background regions is effectively suppressed. During training, the attention weights are automatically adjusted as the loss function backpropagates, enabling the model to gradually develop the ability to explicitly focus on key defect regions, thereby improving the accuracy and interpretability of the overall diagnostic results.

9. The intelligent diagnosis and reliable tracing method for weld defects using X-ray images based on multimodal heterogeneous information collaboration as described in claim 1, characterized in that: The Weld Radiographic Film Defect Dataset (WRF-DD) uses industrial radiographic films collected during actual weld inspections in enterprises. These films are digitized using professional scanning equipment to form high-resolution weld image data. The overall data has significant characteristics such as a real industrial background, complex imaging conditions, and diverse defect morphologies. The data acquisition process strictly follows current non-destructive testing standards for welds and can objectively reflect the imaging characteristics of weld defects in actual engineering inspection environments. In terms of annotation methods, this dataset adopts pixel-level semantic segmentation annotation, in which professionals with experience in industrial non-destructive testing perform pixel-by-pixel fine annotation of the weld defect area according to relevant testing standards to ensure the accuracy and consistency of the defect area outline; semantic segmentation annotation not only includes the spatial location of the defect, but also accurately describes the morphological structure of the defect, which is conducive to the model learning the real geometric features and imaging patterns of weld defects.

10. The intelligent diagnosis and reliable tracing method for weld defects using X-ray images based on multimodal heterogeneous information collaboration as described in claim 1, characterized in that: Compatible automatic digital acquisition and observation equipment for weld X-ray film: The system includes a viewing area (1), with a film delivery area (2) on one side and a film receiving area (3) on the other side. A light shield (101) is provided above the viewing surface of the viewing area (1). Inside the light shield (101), pulleys I (107), II (108), III (109), and IV (110) are installed in sequence via a horizontal belt (111). A motor (102) is installed on the light shield (101) near the film delivery area (2). The motor (102) is used to drive a first longitudinal shaft (103) located near the film delivery area (2). The pulley I (107) is installed on the first longitudinal shaft. (103) At the upper end, the first longitudinal shaft (103) is equipped with an upper feed rubber roller (112) and a lower feed rubber roller (113) at its lower end. The upper feed rubber roller (112) and the lower feed rubber roller (113) are used to press the upper and lower edges of the weld X-ray film. The pulley II (108) and the upper active guide roller I (114) are coaxially mounted on the upper and lower ends of the second longitudinal shaft (104). A lower driven guide roller I (115) is provided below the upper active guide roller I (114). The pulley III (109) and the upper active guide roller II (116) are coaxially mounted on the third longitudinal shaft (104). 5) At the upper and lower ends, a lower driven guide roller II (117) is provided below the upper active guide roller II (116); the pulley IV (110) and the upper receiving rubber roller (118) and the lower receiving rubber roller (119) are coaxially mounted on the fourth longitudinal axis (106); the upper receiving rubber roller (112), the upper active guide roller I (114), the upper active guide roller II (116), and the upper receiving rubber roller (118) are located on the same horizontal straight line, and the lower receiving rubber roller (113), the lower driven guide roller I (115), the lower driven guide roller II (117), and the lower receiving rubber roller (118) are located on the same horizontal straight line. 9) Located on the same horizontal straight line; two sensors (120) are set diagonally in the observation area of ​​the observation area (1); the bottom surface of the observation area of ​​the observation area (1) and the bottom surface of the receiving area (3) are provided with a film guide groove (121) that is connected to the film guide groove (121) and the front baffle (201); the front end of the bottom plate of the film delivery area (2) is fixedly provided with a front baffle (201) and the rear end of the bottom plate is fixedly provided with a rear baffle (202); a movable baffle (203) is provided between the front baffle (201) and the rear baffle (202); a compression spring (204) is connected between the movable baffle (203) and the rear baffle (202). Overall, this weld X-ray semantic segmentation dataset is characterized by its real industrial origin, complex defect morphology, strong background interference, high annotation accuracy, and uneven category distribution. It can effectively support the research and verification of methods related to intelligent segmentation and quality assessment of weld defects.

Citation Information

Cited By

  • Toy bag plastic layer online quality detection method based on multi-modal visual perception

    CN122156213A

  • Online Quality Inspection Method for Toy Plastic Coating Based on Multimodal Visual Perception

    CN122156213B