Nuclear power equipment weld defect detection method based on DBCAM-Trans model
Through the weld defect detection method based on the DBCAM-Trans model, the problems of time-consuming and labor-intensive and subjective factors are solved, and the rapid and accurate detection of weld defects in nuclear power equipment are achieved, which improves quality inspection efficiency and reduces production costs.
Patent Information
- Application Number
- CN202510329882.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
AI Technical Summary
The traditional weld detection method relies on X-ray image acquisition and manual visual inspection, which is time-consuming and labor-intensive, has limited efficiency, and is easily affected by subjective factors. It is impossible to achieve rapid and accurate detection of weld defects in nuclear power equipment, increasing production costs.
Weld defect detection method based on the DBCAM-Trans model is adopted, and the global encoder and local encoder branches are constructed by pre-processing and annotating the X-ray image, combining the multi-scale dual-branch feature fusion module and the global local cross-attention module, the model is trained using cross entropy and FocalLoss, and data enhancement is applied to the Mixup strategy, ultimately achieving rapid and accurate detection of weld defects.
It realizes rapid and accurate detection of weld defects in nuclear power equipment, improves quality inspection efficiency, reduces production costs, improves intelligence level, and can accurately divide weld defect areas, which is suitable for industrial production environments.
Smart Images

Figure CN120259230A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of weld defect detection, and particularly relates to a method for detecting weld defects of nuclear power equipment based on the DBCAM-Trans model. Background Art
[0002] As important equipment in the industrial system, nuclear power equipment and pressure vessels, their safe operation directly affects the safety of industrial production. Although strict controls are taken during production, inspection and acceptance, and installation welding, internal defects still inevitably exist in the material interior and welded joints. These defects not only affect the structural strength of the welded parts, but may also lead to a decline in the performance of the manufactured products, and even pose potential safety hazards. Therefore, timely detection of weld defects existing in industrial equipment plays an important role in the stable operation of the equipment and the safety guarantee of industrial production.
[0003] Traditional weld detection methods mainly rely on X-ray image acquisition and manual visual inspection, which are time-consuming, laborious, limited in efficiency, increase production costs, and are easily affected by subjective factors. Therefore, applying a method for detecting weld defects of nuclear power equipment based on deep learning to the production environment can achieve more accurate, rapid, and cost-effective detection of weld defects in X-ray images, improve defect detection efficiency, and reduce production costs. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for detecting weld defects of nuclear power equipment based on the DBCAM-Trans model, which can achieve rapid and accurate detection of weld defects of nuclear power equipment, improve the quality inspection efficiency, and reduce production costs.
[0005] The technical solutions adopted by the present invention are specifically as follows:
[0006] A method for detecting weld defects of nuclear power equipment based on the DBCAM-Trans model, comprising the following steps:
[0007] Step 1: Preprocess the scanned X-ray image, and the preprocessing includes performing numerical mapping on the X-ray image and adjusting the contrast and brightness of the X-ray image;
[0008] Step 2: Mark the weld defect area in the preprocessed X-ray image, generate corresponding labels according to the marking results, and randomly allocate the X-ray image and the corresponding label image into a training set and a validation set in proportion to form an X-ray image weld defect detection data set;
[0009] Step 3: Construct a DBCAM-Trans model;
[0010] The DBCAM-Trans model includes a global encoder branch, a local encoder branch, and a decoder. The global encoder branch and the local encoder branch both include a multi-scale double-branch feature fusion module, an overlapping image feature sequence extraction layer arranged in a stack, and a global-local cross-attention module. Moreover, both the global encoder branch and the local encoder branch include a multi-scale double-branch feature fusion module, which is used to aggregate the features of each model layer in the global encoder branch and the local encoder branch.
[0011] The outputs of the global encoder branch and the local encoder branch are used as the inputs of the decoder. The decoder continuously reconstructs the image information to obtain the segmentation result of the weld defect area in the X-ray image.
[0012] Step 4: Use the X-ray image weld defect detection dataset constructed in Step 2 to train the DBCAM-Trans model constructed in Step 3, and finally retain the model weights with the best performance on the validation set for use during testing.
[0013] Step 5: Apply the trained DBCAM-Trans model to the X-ray image weld defect detection task.
[0014] Furthermore, Step 1 includes the following steps:
[0015] Step 101: First, perform numerical mapping on the X-ray image I 1 . Using numerical mapping, map the original 16-bit depth X-ray image to an 8-bit depth X-ray image, and map the pixel value range from the original 0 to 65536 to the range of 0 to 255:
[0016]
[0017] where I 1 is the original 16-bit depth X-ray flaw detection image, I 2 is the converted 8-bit depth X-ray image, min(·) is the operation of taking the minimum value of the image, and max(·) is the operation of taking the maximum value of the image.
[0018] Step 102: Adjust the contrast and brightness of the numerically mapped image I 2 . The contrast and brightness adjustment process processes the image according to the following formula:
[0019] I 3 = αI 2 + β
[0020] where I 2 represents the numerically mapped X-ray image, and I 3It is the adjusted X-ray image, where α is the coefficient for adjusting the image contrast, α > 0, and β is the gain variable.
[0021] Furthermore, the global encoder branch in step 3 includes a stacked overlapping image feature sequence extraction layer and a global-local cross-attention module. The overlapping image feature sequence extraction layer is located at the first layer of the global encoder branch, followed by 12 consecutively stacked global-local cross-attention modules. Each global-local cross-attention module has 12 heads and a feature dimension of 768.
[0022] The local encoder branch in step 3 includes a stacked overlapping image feature sequence extraction layer and a global-local cross-attention module. The overlapping image feature sequence is located at the first layer of the local encoder branch, followed by 12 consecutively stacked global-local cross-attention modules. Each global-local cross-attention module has 6 heads and a feature dimension of 384.
[0023] Furthermore, the local encoder branch randomly crops a sub-region image from the complete X-ray image as input data. The central position of the sub-region is randomly selected within the 30% region of the image center, and both the width and height are selected within the range of 10% to 90% of the width and height of the complete X-ray image. The cropped sub-region image is scaled to the same input size as the global encoder branch when input into the local encoder branch.
[0024] Furthermore, the overlapping feature sequence extraction layer includes two layers of conventional convolutional layers, a single layer of batch normalization layer, and a single layer of ReLU activation layer. The convolutional kernel size of the conventional convolutional layer is 17, the stride is 3, and the padding number is 3. The batch normalization layer and the ReLU activation layer are located between the two convolutional layers, and the ReLU activation layer is after the batch normalization layer. After the X-ray image is convolved by the overlapping feature sequence extraction layer and flattened into one dimension in the first two dimensions, it becomes the initial feature sequence.
[0025] Among them, the initial feature sequence of the global encoder branch is The initial feature sequence of the local encoder branch is Among them, N and D represent the number of feature blocks and the feature dimension of a single feature block of the feature sequence, respectively.
[0026] Furthermore, the global-local cross-attention module includes a sequentially connected global-local cross-self-attention layer, a first layer normalization layer, a conventional multi-head self-attention layer, a multi-layer perceptron, and a second layer normalization layer. The output features of the first layer normalization layer are added to the input features of the global-local cross-attention module, and the output features of the second layer normalization layer are added to the output features of the first layer normalization layer to form two sets of skip connections.
[0027] The global-local cross self-attention layer includes query sequence, key sequence, and value sequence obtained by mapping through three mutually independent linear layers. The query sequence and the key sequence are used to calculate the attention weights. After scaling in the feature dimension, the value sequence is weighted, and the weighted result is used as the output of the global-local cross self-attention layer;
[0028] The conventional multi-head self-attention layer also includes these three groups of feature sequences: query sequence, key sequence, and value sequence, which are obtained by mapping through three mutually independent linear layers. After calculating the attention weights between the query sequence and the key sequence, the value sequence is weighted to obtain the output of the conventional multi-head self-attention layer. The three groups of feature sequences are all mapped from the features of the branch where the global-local cross-attention module is located.
[0029] Further, the multi-scale double-branch feature fusion module in step 3 includes a linear layer. The features of the 3rd, 6th, 9th, and 12th layers of the global encoder branch and the local encoder branch are respectively taken. The features from the two branches at each depth are concatenated in the feature dimension. A single-layer linear layer is used to reduce the dimension of the feature fusion. Thus, the features fused by the linear layer at four depths are obtained. The four groups of fused features are concatenated again in the feature dimension, and then the features are fused and dimension-reduced again through a multi-layer perceptron. Finally, the features pass through a layer normalization layer and a Dropout layer with a dropout rate of 50% to become the encoder output features after the fusion of the global encoder branch and the local encoder branch.
[0030] Further, the decoder includes three groups of decoder modules and a linear layer. The decoder module includes a transposed convolution layer, a batch normalization layer, and an activation layer. The decoder restores the encoder output features into a feature map, and then sequentially passes through three groups of decoder modules, gradually restoring the information through the transposed convolution layer from the low-resolution feature map. After passing through the linear layer with an output dimension of 2 and the SoftMax activation layer at the last layer, it becomes the probability indicating that each pixel position belongs to the defective area.
[0031] Further, when training the DBCAM-Trans model in step 4, the cross-entropy and FocalLoss are used to train the DBCAM-Trans model. The cross-entropy and FocalLoss loss functions are fused according to the specific weight coefficients α and β according to
[0032] α is usually set to 2 and β is set to 1;
[0033] In step 4, the Mixup strategy is used to perform data augmentation on the input images during the training of the DBCAM-Trans model. Mixup mixes the input images and labels of two samples in a certain proportion to improve the generalization ability of the model by blurring the boundaries between samples. The mixing process is expressed as:
[0034] I mix = mIa +nI b
[0035] Y mix = mY a +nY b
[0036] where I is the input image, Y is the label, and I mix and Y mix are the mixed input image and the mixed label respectively, and m and n are the weight coefficients used when mixing two input images and two labels.
[0037] Furthermore, when actually applying the trained DBCAM-Trans model to the X-ray image weld defect detection task in step 5, the optimal model weights are pre-loaded, the X-ray image to be detected is pre-processed, and then input into the DBCAM-Trans model loaded with the optimal model weights to obtain a binary image indicating the position of the weld defect.
[0038] The technical effects achieved by the present invention are as follows:
[0039] A method for detecting weld defects in nuclear power equipment based on the DBCAM-Trans model of the present invention is based on high-quality weld defect image data collected and labeled in real time, constructs and trains a method for detecting weld defects in nuclear power equipment based on the DBCAM-Trans model, accurately segments the weld defect area in nuclear power equipment, realizes rapid and accurate detection of weld defects in nuclear power equipment, can be applied to the weld defect detection task in the industrial production environment, realizes rapid detection of the weld defect area by analyzing the industrial equipment scan image, improves the quality inspection efficiency, enhances the intelligent level, reduces the production cost. The advantage of the DBCAM-Trans model is that it can fully combine the global information and local details of the image to predict the position and shape of the weld defect area, and can achieve a more accurate weld area segmentation result. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is the training and application flowchart of the present invention;
[0041] Figure 2 is the network structure diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0042] In order to make the purpose and advantages of the present invention clearer, the present invention will be specifically described below in conjunction with the embodiments. It should be understood that the following text is only used to describe one or several specific implementation manners of the present invention, and does not strictly limit the specific protection scope claimed by the present invention.
[0043] First, the present invention improves the quality of the collected X-ray images through image preprocessing, and constructs a training set and a validation set after annotation by professionals. Then, the DBCAM-Trans model is trained on the constructed data set. Finally, when applied in practice, the trained model is used to extract the rich features contained in the X-ray images, obtaining an accurate and clear segmentation result of the defect area, providing assistance for the rapid detection of weld defects.
[0044] As Figure 1-2 shown, a method for detecting weld defects in nuclear power equipment based on the DBCAM-Trans model includes the following steps:
[0045] Step 1: Preprocess the scanned X-ray images;
[0046] Specifically, Step 1 includes the following steps:
[0047] Step 101: First, perform numerical mapping on the X-ray image I 1 . Using numerical mapping, map the original 16-bit depth X-ray image to an 8-bit depth X-ray image, and map the pixel value range from the original 0 to 65536 to the range of 0 to 255:
[0048]
[0049] where, I 1 is the original 16-bit depth X-ray flaw detection image, I 2 is the converted 8-bit depth X-ray flaw detection image, that is, the X-ray image after numerical mapping, min(·) is the operation of taking the minimum value of the image, and max(·) is the operation of taking the maximum value of the image.
[0050] Step 102: Adjust the contrast and brightness of the numerically mapped image I 2 . There are a series of problems in the original X-ray images, such as low contrast, small gray values, the overall image being too dark, and the gray distribution being concentrated. Therefore, adjusting its brightness and contrast can effectively improve the distribution of pixel values in each gray range and enhance the recognition ability of image features. The process of adjusting contrast and brightness is to process the image according to the following formula:
[0051] I 3 = αI 2 + β
[0052] where, I 2 represents the X-ray image after numerical mapping, I 3 is the adjusted X-ray image, α is the coefficient for adjusting the contrast of the image, α > 0, and β is the gain variable. α can make the image pixels increase or decrease multiplicatively, changing the contrast of the image. β is to directly superimpose a change amount on the original pixel value, changing the brightness of the image;
[0053] The contrast and brightness adjustment in step 102 is performed according to the distribution of image pixel values in each gray scale range, so as to improve the recognition ability of image features.
[0054] Step 2: A professional manually marks the weld defect areas in the preprocessed X-ray image I 3 , generates a corresponding label image Y based on the results of the manual marking, and randomly assigns all X-ray flaw detection images and the corresponding label images to the training set (X train , Y train ) and the validation set (X eval , Y eval ) in a ratio of 8:2, forming an X-ray image weld defect detection data set.
[0055] Step 3: Construct a DBCAM-Trans model;
[0056] The full name of the DBCAM-Trans model is the Dual-Branch Cross-Attention Multiscale Transformer model based on multi-scale cross-attention.
[0057] As Figure 2 shown, the DBCAM-Trans model constructed in the present invention mainly includes three parts: a global encoder branch, a local encoder branch, and a decoder. Among them, both the global encoder branch and the local encoder branch include an overlapping image feature sequence extraction layer and a global-local cross-attention module stacked, and both the global encoder branch and the local encoder branch are stacked by an overlapping image feature sequence extraction layer and a global-local cross-attention module, and both the global encoder branch and the local encoder branch include a multi-scale dual-branch feature fusion module, and the multi-scale dual-branch feature fusion module is used to aggregate the features of each model layer in the global encoder branch and the local encoder branch.
[0058] The outputs of the global encoder branch and the local encoder branch are used as the inputs of the decoder, and the decoder continuously reconstructs the image information through deconvolution layers to obtain the segmentation result of the weld defect area in the X-ray image. The advantage of the global-local cross-attention module compared with the prior art is that it enables the model to pay attention to both global information and local detail information while, making the predicted defect area segmentation result edge clearer and more accurate. The advantage of the multi-scale dual-branch feature fusion module compared with the prior art is that it enables the model to fuse global information and local details at each image scale, ensuring that the global information and local details can cooperate effectively.
[0059] Among them, the global encoder branch in step 3 includes an overlapping image feature sequence extraction layer and a global-local cross-attention module arranged in a stack. The global encoder branch is stacked by the overlapping image feature sequence extraction layer and the global-local cross-attention module. The overlapping image feature sequence extraction layer is located at the first layer of the global encoder branch, followed by 12 continuously stacked global-local cross-attention modules. Each global-local cross-attention module has 12 heads and a feature dimension of 768. The global encoder branch processes the complete X-ray image and focuses on extracting the global features of the X-ray image through a larger number of heads and a larger feature dimension.
[0060] Among them, the local encoder branch in step 3 includes an overlapping image feature sequence extraction layer and a global-local cross-attention module arranged in a stack. The local encoder branch is also stacked by the overlapping image feature sequence extraction layer and the global-local cross-attention module. The overlapping image feature sequence is also located at the first layer of the local encoder branch, followed by 12 continuously stacked global-local cross-attention modules. Each global-local cross-attention module has 6 heads and a feature dimension of 384. The local encoder branch processes the sub-region of the X-ray image, that is, randomly cropping the sub-region image from the complete X-ray image as the input data. Among them, the central position of the sub-region is randomly selected within the area of 30% of the image center, and the width and height are both selected within the range of 10% to 90% of the width and height of the complete X-ray image. All the cropped sub-region images are scaled to the same input size as the global encoder branch when input into the local encoder branch. The local encoder branch focuses on extracting the local features of the X-ray image through a smaller input scale and comparable model parameters.
[0061] Among them, the overlapping feature sequence extraction layer includes two layers of conventional convolutional layers, a single layer of batch normalization layer, and a single layer of ReLU activation layer. The convolutional kernel size of the conventional convolutional layer is 17, the stride is 3, and the padding number is 3. Among them, the batch normalization layer and the ReLU activation layer are located between the two convolutional layers, and the ReLU activation layer is after the batch normalization layer, that is, the output result of the batch normalization layer is used as the input of the ReLU activation layer. X-ray flaw detection image (X-ray image) I 3 After being convolved by the overlapping feature sequence extraction layer and flattened into one dimension in the first two dimensions, it becomes the initial feature sequence.
[0062] Among them, the initial feature sequence of the global encoder branch is The initial feature sequence of the local encoder branch is Among them, N and D represent the number of feature blocks of the feature sequence and the feature dimension of a single feature block, respectively.
[0063] The X-ray image passes through the overlapping feature sequence extraction layer of the global encoder branch to obtain the initial feature sequence of the global encoder branch where N and D1 represent the sequence length and feature dimension of the initial feature sequence of the global encoder branch, respectively;
[0064] The X-ray image passes through the overlapping feature sequence extraction layer of the local encoder branch to obtain the initial feature sequence of the local encoder branch where N and D2 represent the sequence length and feature dimension of the initial feature sequence of the local encoder branch, respectively.
[0065] Among them, the global-local cross-attention module includes a global-local cross-self-attention layer, a first layer normalization layer, a conventional multi-head self-attention layer, a multi-layer perceptron, and a second layer normalization layer connected in sequence. The global-local cross-attention module is sequentially composed of five model layers: the global-local cross-self-attention layer, the first layer normalization layer, the conventional multi-head self-attention layer, the multi-layer perceptron, and the second layer normalization layer connected in sequence. And the output features of the first layer normalization layer are added to the input features of the global-local cross-attention module, and the output features of the second layer normalization layer are added to the output features of the first layer normalization layer, thus forming two sets of skip connections to alleviate the problem of gradient disappearance during training.
[0066] The global-local cross-self-attention layer includes three groups of feature sequences: a query sequence, a key sequence, and a value sequence, which are obtained by mapping through three mutually independent linear layers. The query sequence calculates the attention weights with the key sequence, and after scaling the feature dimension, weights the value sequence, and the weighted result is used as the output of the global-local cross-self-attention layer.
[0067] The conventional multi-head self-attention layer also includes three groups of feature sequences: a query sequence, a key sequence, and a value sequence, which are obtained by mapping through three mutually independent linear layers. After calculating the attention weights between the query sequence and the key sequence, the value sequence is weighted to obtain the output of the conventional multi-head self-attention layer, where the three groups of feature sequences are all mapped from the features of the branch where the global-local cross-attention module is located.
[0068] Among them, the key of the global-local cross-self-attention layer lies in the query sequence the key sequence the value sequence The feature sources are different from those of the conventional multi-head self-attention layer. The three groups of feature sequences of the conventional multi-head self-attention layer are obtained by mapping the same set of encoder features through three mutually independent linear layers. The source of the query sequence in the global-local cross-self-attention layer is different from the remaining two groups of feature sequences (the key sequence the value sequence ) are from different sources. In the global encoder branch, the query sequence is obtained from the features of the local encoder branch through a linear layer, while the key sequence and the value sequence are both mapped from the features of the global encoder branch. In the local encoder branch, the query sequence is obtained from the features of the global encoder branch through a linear layer, while the key sequence and the value sequence are both mapped from the features of the local encoder branch.
[0069] Among them, the feature sequence of the global encoder branch is sequentially mapped to the query sequence of the local encoder branch the key sequence of the global encoder branch and the value sequence The feature sequence of the local encoder branch is sequentially mapped to the query sequence of the global encoder branch the key sequence of the local encoder branch and the value sequence Then, the global-local cross self-attention layers of the global encoder branch and the local encoder branch respectively perform multi-head self-attention calculations, that is, the feature dimension is equally divided into multiple groups of features, and each group of features is one head. In each head, the query sequence calculates the attention weights with the key sequence, and after scaling in the feature dimension, the value sequence is weighted. After the weighted results of multiple heads are combined, the output of the conventional multi-head self-attention layer is obtained. The process can be expressed as:
[0070]
[0071]
[0072] Among them, T represents matrix transpose, is the output feature of the i-th global-local cross-attention module in the global encoder branch, is the output feature of the i-th global-local cross-attention module in the local encoder branch, and D1 and D2 are the feature dimensions of the global encoder branch and the local encoder branch respectively. and are the output features of the global-local cross self-attention layer in the global-local cross-attention module. Subsequently, the feature sequence is processed by a layer normalization layer and then passes through a conventional multi-head self-attention layer to be further aggregated into features and Among them, the query sequence, key sequence, and value sequence of the conventional multi-head self-attention layer all come from the feature sequence of the corresponding branch. The subsequent multi-layer perceptron contains two linear layers. The first linear layer first expands the feature dimensions of the feature sequences and from D1 and D2 to 4D1 and 4D2, and the second linear layer then reduces its feature dimensions back to D1 and D2. After skip connection, the output feature sequence of the (i + 1)-th global-local cross-attention module is obtained and
[0073] Among them, the multi-scale double-branch feature fusion module in step 3 includes a linear layer, which respectively takes the features of the 3rd, 6th, 9th, and 12th layers of the global encoder branch and the local encoder branch. The features obtained from the global encoder branch and the local encoder branch are respectively denoted as and For the features from the two branches at each depth, they are concatenated in the feature dimension. At this time, the feature dimension is D1+D2, and then the feature fusion is reduced in dimension by a single-layer linear layer, thereby obtaining the features after linear layer fusion at four depths The dimensions of the fused features are all D. The four groups of fused features are concatenated again in the feature dimension, and then the features are fused and reduced in dimension by a multi-layer perceptron. Among them, the input dimension of the multi-layer perceptron module is 4D, the hidden layer dimension is 2D, and the output dimension is D. Finally, the features pass through a layer normalization layer and a Dropout layer with a dropout rate of 50% to become the encoder output features after the fusion of the global encoder branch and the local encoder branch
[0074] As Figure 2 shown, the layer normalization layer and the Dropout layer passed by the above features exist between the multi-layer perceptron and the decoder, and do not belong to other modules, which are used to normalize and randomly discard the features, helping to improve the model training speed and the model generalization ability
[0075] Among them, the decoder in step 3 includes three groups of decoder modules and a linear layer. The decoder module includes a transposed convolution layer, a batch normalization layer, and an activation layer. The decoder restores the encoder output features into a feature map Then it passes through three groups of decoder modules in sequence, gradually mapping from the low-resolution feature map to the segmentation result. The transposed convolution layer in each group of decoder modules maps the feature map to twice the original size through transposed convolution operation to restore information. The batch normalization layer scales the feature values to ensure the smoothness of training. The activation layer is usually a ReLU non-linear mapping, which can reduce the vanishing gradient and improve the convergence speed. The linear layer is the last model layer of the decoder. The output dimension of the linear layer of the decoder is 2. After the feature map is mapped by the linear layer and then activated by SoftMax, it becomes the probability indicating that each pixel position belongs to the defect area
[0076] Step 4: Train the DBCAM-Trans model constructed in step 3 with the X-ray image weld defect detection dataset constructed in step 2
[0077] During training, Mixup is used to augment the X-ray image weld defect detection dataset used as training data. Cross-entropy and Focal Loss are used to calculate the difference (loss function) between the output of the DBCAM-Trans model and the target output. Finally, the model weights with the best performance on the validation set are retained for use during testing.
[0078] The DBCAM-Trans model is trained using cross-entropy and Focal Loss. Cross-entropy is the most commonly used basic constraint for classification and segmentation problems, which can ensure that the model converges as soon as possible in a short time. Focal Loss is suitable for situations where the target area accounts for a relatively small proportion, such as defect area segmentation, and there is a significant imbalance in the number of category samples, which can improve the final accuracy of the model. The final loss function adopted is the combination of cross-entropy and Focal Loss, that is:
[0079]
[0080] Among them, is the loss function, and represent cross-entropy and Focal Loss respectively, and α and β are the weight coefficients when adding the two losses. Usually, α is set to 2 and β is set to 1.
[0081] In step 4, the Mixup strategy is used to perform data augmentation on the input images during the training of the DBCAM-Trans model. Mixup mixes the input images and labels of two samples in a certain proportion to improve the generalization ability of the model by blurring the boundaries between samples. The mixing process can be expressed as:
[0082] I mix = mI a + nI b
[0083] Y mix = mY a + nY b
[0084] Among them, I is the input image, Y is the label, I mix and Y mix are the mixed input image and the mixed label respectively, and m and n are the weight coefficients used when mixing two input images and two labels.
[0085] Step 5: When applying the trained DBCAM-Trans model to the X-ray image weld defect detection task, first use the same preprocessing method as for processing the training data to preprocess the X-ray image to be analyzed, including numerical mapping and contrast and brightness adjustment. Then input the preprocessed image into the trained DBCAM-Trans model to obtain the probability distribution map of the weld defect area.
[0086] Among them, the probability map of the weld defect area output by the decoder contains two channels. For each pixel position, if the probability value of the first channel is larger, the category of this pixel position is the non-defect area; if the probability value of the second channel is larger, the category of this pixel position is the defect area. Thus, a binary map indicating the defect area is obtained. This binary map will be used to assist the weld defect detection task of professional defect detection personnel in the actual production environment.
[0087] Moreover, when actually applying the X-ray image weld defect area detection, preload the optimal model weights, preprocess the X-ray image to be detected, and input it into the DBCAM-Trans model loaded with the optimal model weights to obtain a binary image indicating the weld defect position, providing assistance for defect detection personnel.
[0088] In summary, based on the high-quality weld defect image data collected and labeled in reality, this technical solution constructs and trains a nuclear power equipment weld defect detection method based on the DBCAM-Trans model, accurately segments the weld defect area of nuclear power equipment, and realizes the rapid and accurate detection of nuclear power equipment weld defects. The method of the present invention can be applied to the weld defect detection task in the industrial production environment. By analyzing the industrial equipment scanning images, it can realize the rapid detection of the weld defect area, improve the quality inspection efficiency, enhance the intelligent level, and reduce the production cost. The advantage of the DBCAM-Trans model is that it can fully combine the global information and local details of the image to predict the position and shape of the weld defect area, and can achieve a more accurate weld area segmentation result.
[0089] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. The structures, devices, and operation methods not specifically described and explained in the present invention are implemented according to the conventional means in the art without special description and limitation.
Claims
1. A method for detecting weld defects in nuclear power equipment based on the DBCAM-Trans model, characterized in that: It includes the following steps: Step 1: Preprocess the scanned X-ray images, including numerical mapping of the X-ray images and adjustment of the contrast and brightness of the X-ray images; Step 2: Annotate the weld defect areas in the preprocessed X-ray images, generate corresponding labels according to the annotation results, and randomly allocate the X-ray images and the corresponding label images into a training set and a validation set in proportion to form an X-ray image weld defect detection data set; Step 3: Construct a DBCAM-Trans model; The DBCAM-Trans model includes a global encoder branch, a local encoder branch, and a decoder; Among them, both the global encoder branch and the local encoder branch include a multi-scale double-branch feature fusion module and a stacked overlapping image feature sequence extraction layer and a global-local cross-attention module, and both the global encoder branch and the local encoder branch include a multi-scale double-branch feature fusion module. The multi-scale double-branch feature fusion module is used to aggregate the features of each model layer in the global encoder branch and the local encoder branch; The outputs of the global encoder branch and the local encoder branch are used as the inputs of the decoder, and the decoder continuously reconstructs the image information to obtain the segmentation result of the weld defect area in the X-ray image; Step 4: Train the DBCAM-Trans model constructed in Step 3 with the X-ray image weld defect detection data set constructed in Step 2, and finally retain the model weights with the best performance on the validation set for use during testing; Step 5: Apply the trained DBCAM-Trans model to the X-ray image weld defect detection task.
2. The method for detecting weld defects of nuclear power equipment based on the DBCAM-Trans model according to claim 1, wherein: Step 1 includes the following steps: Step 101: First, perform numerical mapping on the X-ray image I 1 Perform numerical mapping to map the original 16-bit depth X-ray image to an 8-bit depth X-ray image, and map the pixel value range from the original 0 to 65536 to the range of 0 to 255: Among them, I 1 is the original X-ray flaw detection image with 16-bit depth, and I 2 is the converted X-ray image with 8-bit depth. min(·) is the operation of taking the minimum value of the image, and max(·) is the operation of taking the maximum value of the image; Step 102: For the image I after numerical mapping 2 Perform contrast and brightness adjustment. The contrast and brightness adjustment process processes the image according to the following formula: I 3 = αI 2 + β Among them, I 2 represents the X-ray image after numerical mapping, and I 3 is the adjusted X-ray image, α is the coefficient for adjusting the image contrast, α > 0, and β is the gain variable.
3. A method for detecting weld defects of nuclear power equipment based on the DBCAM-Trans model according to claim 1, characterized in that: The global encoder branch in Step 3 includes a stacked overlapping image feature sequence extraction layer and a global-local cross-attention module. The overlapping image feature sequence extraction layer is located at the first layer of the global encoder branch, followed by 12 continuously stacked global-local cross-attention modules. The number of heads of each global-local cross-attention module is 12, and the feature dimension is 768; The local encoder branch in Step 3 includes a stacked overlapping image feature sequence extraction layer and a global-local cross-attention module. The overlapping image feature sequence is located at the first layer of the local encoder branch, followed by 12 continuously stacked global-local cross-attention modules. The number of heads of each global-local cross-attention module is 6, and the feature dimension is 384.
4. The method for detecting weld defects of nuclear power equipment based on the DBCAM-Trans model according to claim 3, characterized in that: The local encoder branch randomly crops sub-region images from the complete X-ray image as input data. The central position of the sub-region is randomly selected within the 30% area of the image center, and both the width and height are selected within the range of 10% to 90% of the width and height of the complete X-ray image. The cropped sub-region images are scaled to the same input size as the global encoder branch when input into the local encoder branch.
5. The method for detecting weld defects of nuclear power equipment based on the DBCAM-Trans model according to claim 3, wherein: The overlapping feature sequence extraction layer includes two layers of conventional convolutional layers, one layer of batch normalization layer, and one layer of ReLU activation layer. The convolutional kernel size of the conventional convolutional layer is 17, the stride is 3, and the padding number is 3. The batch normalization layer and the ReLU activation layer are located between the two convolutional layers, and the ReLU activation layer is after the batch normalization layer. After the X-ray image is convolved by the overlapping feature sequence extraction layer and flattened into one dimension in the first two dimensions, it becomes the initial feature sequence. Among them, the initial feature sequence of the global encoder branch is The initial feature sequence of the local encoder branch is Among them, N and D represent the number of feature blocks of the feature sequence and the feature dimension of a single feature block, respectively.
6. A method for detecting weld defects of nuclear power equipment based on the DBCAM-Trans model according to claim 3, characterized in that: The global-local cross-attention module includes a global-local cross-self-attention layer, a first layer normalization layer, a conventional multi-head self-attention layer, a multi-layer perceptron, and a second layer normalization layer connected in sequence. The output features of the first layer normalization layer are added to the input features of the global-local cross-attention module, and the output features of the second layer normalization layer are added to the output features of the first layer normalization layer to form two sets of skip connections. The global-local cross-self-attention layer includes a query sequence, a key sequence, and a value sequence obtained by mapping through three mutually independent linear layers. The query sequence calculates the attention weights with the key sequence, and after feature dimension scaling, the value sequence is weighted, and the weighted result is used as the output of the global-local cross-self-attention layer. The conventional multi-head self-attention layer also includes three sets of feature sequences, namely a query sequence, a key sequence, and a value sequence, obtained by mapping through three mutually independent linear layers. After calculating the attention weights between the query sequence and the key sequence, the value sequence is weighted to obtain the output of the conventional multi-head self-attention layer. The three sets of feature sequences are all mapped from the features of the branch where the global-local cross-attention module is located.
7. A method for detecting weld defects of nuclear power equipment based on the DBCAM-Trans model according to claim 6, characterized in that: The multi-scale double-branch feature fusion module in step 3 includes a linear layer. The features of the 3rd, 6th, 9th, and 12th layers of the global encoder branch and the local encoder branch are respectively taken. The features from the two branches at each depth are concatenated in the feature dimension, and a single linear layer is used to reduce the dimension of the feature fusion, thereby obtaining the features after linear layer fusion at four depths. The four sets of fused features are concatenated again in the feature dimension, and then the features are fused and reduced in dimension again through a multi-layer perceptron. Finally, the features pass through a layer normalization layer and a Dropout layer with a dropout rate of 50% to become the encoder output features after the fusion of the global encoder branch and the local encoder branch.
8. A method for detecting weld defects of nuclear power equipment based on the DBCAM-Trans model according to claim 1, characterized in that: The decoder includes three sets of decoder modules and a linear layer. The decoder module includes a transposed convolutional layer, a batch normalization layer, and an activation layer. The decoder restores the encoder output features into a feature map, and then sequentially passes through three sets of decoder modules, gradually restoring information through the transposed convolutional layer from the low-resolution feature map. After passing through the linear layer with an output dimension of 2 and the SoftMax activation layer at the last layer, it becomes the probability indicating that each pixel position belongs to the defect area.
9. A method for detecting weld defects of nuclear power equipment based on the DBCAM-Trans model according to claim 1, characterized in that: When training the DBCAM-Trans model in step 4, the cross-entropy and FocalLoss are used to train the DBCAM-Trans model. The cross-entropy and FocalLoss loss functions are fused according to specific weight coefficients α and β. Usually, α is set to 2 and β is set to 1; In step 4, the Mixup strategy is used to perform data augmentation on the input images during the training of the DBCAM-Trans model. Mixup mixes the input images and labels of two samples in a certain proportion to improve the generalization ability of the model by blurring the boundaries between samples. The mixing process is expressed as: I mix = mI a + nI b Y mix = mY a + nY b Where, I is the input image, Y is the label, I mix and Y mix are the mixed input image and the mixed label respectively, and m and n are the weight coefficients used when mixing two input images and two labels.
10. A method for detecting weld defects of nuclear power equipment based on the DBCAM-Trans model according to claim 1, characterized in that: When the trained DBCAM-Trans model is actually applied to the X-ray image weld defect detection task in step 5, the optimal model weights are pre-loaded, the X-ray image to be detected is pre-processed, and it is input into the DBCAM-Trans model loaded with the optimal model weights to obtain a binary image indicating the position of the weld defect.