Nuclear power pipeline X-ray flaw detection image defect detection method based on Transform distillation model

Through the lightweight image processing method based on the Transformer distillation model and the experience of professional film reviewers to build the training set, the problem of manual judgment workload in X-ray flaw detection image defect detection in nuclear power pipelines is solved, and fast and accurate defect detection is achieved, reducing calculation and training costs.

CN120339186APending Publication Date: 2025-07-18烟台市标准计量检验检测中心(国家蒸汽流量计量烟台检定站烟台市质量技术监督评估鉴定所)
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510315796.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, the detection of X-ray flaw detection image defects in nuclear power pipelines relies on manual judgment, resulting in large workload, high cost and low accuracy.

Method used

A lightweight image processing method based on Transformer distillation model is adopted, and the training set is constructed based on the experience of professional reviewers. The defect area is segmented through the encoder and decoder, and iterative updates and loss calculations are used to reduce the calculation amount and training cost.

Benefits of technology

It realizes fast and accurate defect detection, reduces calculation and training costs, improves detection efficiency, and saves production time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339186A_ABST
    Figure CN120339186A_ABST
Patent Text Reader

Abstract

The invention discloses a nuclear power pipeline X-ray flaw detection image defect detection method based on a Transform distillation model. The nuclear power pipeline X-ray flaw detection image defect detection method comprises the following steps: step 1, preprocessing an obtained scanned nuclear power pipeline X-ray flaw detection image; step 2, a Transform distillation model is constructed; 3, constructing a teacher model branch; step 4, training a Transform distillation model and a teacher model branch; and step 5, detecting the preprocessed X-ray flaw detection image to be detected by using the trained model. The method has the beneficial effects that a deep learning model is provided, nuclear power pipeline X-ray flaw detection image defect detection based on the lightweight Transform distillation model is realized, and an effective auxiliary means is provided for evaluation personnel to realize rapid classification and evaluation of X-ray flaw detection image defect areas.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and specifically to a method for defect detection of X-ray flaw detection images of nuclear power pipelines based on a Transformer distillation model. Background Art

[0002] Nuclear energy is an economic, safe, reliable and clean energy source. Nuclear power converts nuclear energy into electrical energy, which is a common way to utilize nuclear energy. Safety is an important restrictive factor for the development of nuclear power. Although strict controls are taken during the production, inspection and acceptance, and installation and welding of nuclear power plants, it is still inevitable that there are internal defects in the material and the welded connection parts, which will gradually and slowly germinate, expand and grow, gradually form surface and through cracks, and eventually rupture. This will seriously threaten the safety of surrounding building structures, nuclear power safety equipment and staff, and even bring serious secondary disasters such as nuclear leakage. Therefore, the research on flaw detection of nuclear island equipment products is very necessary.

[0003] Currently, the commonly used flaw detection method is X-ray flaw detection. During ray detection, it is necessary to manually determine the position and shape of the defect area in each X-ray flaw detection image. From the initial evaluation to the re-evaluation, the manual detection method makes the workload of relevant film evaluation personnel extremely large. It not only consumes a large amount of manpower, material and financial resources, increases the time cost of production, but also has the problem of low accuracy. Summary of the Invention

[0004] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a deep learning model to realize the defect detection of X-ray flaw detection images of nuclear power pipelines based on a lightweight Transformer distillation model, and provide an effective auxiliary means for film evaluation personnel to quickly classify and evaluate the defect areas in X-ray flaw detection images.

[0005] The purpose of the present invention is achieved by the following technical measures: A method for defect detection of X-ray flaw detection images of nuclear power pipelines based on a Transformer distillation model, including the following steps:

[0006] Step 1: Preprocess the obtained scanned X-ray flaw detection images of nuclear power pipelines, construct a training set and a validation set, and obtain their corresponding mask label sets based on the training set and the validation set;

[0007] Step 2: Construct a Transformer distillation model to segment the defect regions of X-ray flaw detection images. The Transformer distillation model includes an encoder and a decoder. The encoder includes an image patch feature extraction module and a sequence compression multi-head self-attention module. The primary features of the X-ray flaw detection image are extracted by the image patch feature extraction module, and the primary features of the X-ray flaw detection image are enhanced by the sequence compression multi-head self-attention module to obtain output features. The decoder is used to obtain the prediction result of the defect region of the X-ray flaw detection image based on the output features extracted by the encoder.

[0008] Step 3: Construct a teacher model branch, which includes an image patch feature extraction module and a multi-head self-attention module. The primary features are extracted by the image patch feature extraction module, and the output features are obtained by the multi-head self-attention module.

[0009] Step 4: Train the Transformer distillation model and the teacher model branch based on the training set, validation set, and their corresponding mask label sets. During training, the parameters of the teacher model branch are not updated by backpropagation of gradients, but are iteratively updated based on the parameters of the encoder. Calculate the mean squared error loss based on the output features of the teacher model branch and the encoder, and use the sum of the mean squared error loss and the exponentially weighted Focal Loss as the total loss. Use the gradient descent method to perform backpropagation of gradients and update the parameters of the Transformer distillation model based on the total loss.

[0010] Step 5: After training the Transformer distillation model, use the trained model to detect the preprocessed X-ray flaw detection image to be detected.

[0011] In some embodiments, the image patch feature extraction module includes three convolutional layers, two batch normalization layers, and two ReLU activation layers. The ReLU activation layer follows the batch normalization layer, and the batch normalization layer and the ReLU activation layer are located between adjacent convolutional layers.

[0012] In some embodiments, the output features of the third convolutional layer are added to the output features of the first ReLU activation layer to form a set of skip connections.

[0013] In some embodiments, the sequence compression multi-head self-attention module includes a sequence compression multi-head self-attention layer, a first normalization layer, a multi-layer perceptron, and a second normalization layer connected in sequence.

[0014] In some embodiments, the sequence-compressed multi-head self-attention layer first uses three linear layers to map the output features of the image patch feature extraction module into a query sequence, a key sequence, and a value sequence. The key sequence and the value sequence are each passed through a convolutional layer for sequence compression, and then multi-head self-attention calculation is performed based on the query sequence and the compressed key sequence and value sequence.

[0015] In some embodiments, the output features of the first normalization layer are added to the input features of the sequence-compressed multi-head self-attention layer, and the output features of the second normalization layer are added to the output features of the first normalization layer to form two sets of skip connections.

[0016] In some embodiments, the decoder includes a decoder module, and the decoder module includes an upsampling layer and a linear layer. The feature image is enlarged in size through the upsampling layer, and the feature image is dimension-reduced through the linear layer to obtain the final output features with a feature dimension of 2. The final output features are activated by SoftMax to obtain the probability that the current pixel position belongs to the normal or defective area.

[0017] In some embodiments, the feature dimension after dimension reduction by the linear layer of the decoder module corresponds one-to-one with the output dimension after three convolutional layers of the image patch feature extraction module, and the feature map of the image patch feature extraction module is concatenated with the feature map output by the current linear layer to form a long-range skip connection.

[0018] In some embodiments, the iterative calculation process of the teacher model branch is Θ′ t = 0.5(Θ t + Θ s ), where Θ t ′ represents the parameters of the updated teacher model branch, Θ t represents the parameters of the teacher model branch before update, and Θ s represents the parameters of the current encoder.

[0019] In some embodiments, the calculation process of the exponential weighted Focal Loss is where is the probability that the prediction at a single pixel position is the normal area, y comes from the set of mask labels in the training set, i.e., y ∈ Y train , representing the true label at this pixel position, α is the weight coefficient used to alleviate the class imbalance problem, and γ is the exponential coefficient that makes the model pay more attention to hard samples.

[0020] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention proposes a method for defect detection of X-ray flaw detection images of nuclear power pipelines based on a lightweight Transformer distillation model. A training set is constructed by combining the empirical knowledge and film judgment habits of professional film reviewers. An encoder is constructed to extract the feature information of X-ray flaw detection images. The primary features of X-ray flaw detection images are enhanced through a sequence compression multi-head self-attention module, enabling the model to capture deeper and more complex features, while avoiding a significant increase in the computational complexity of the model and ensuring the computational efficiency. The encoder can significantly reduce the computational complexity while ensuring the defect detection effect, and reduce the application cost. The image information is reconstructed through a decoder, and specific defect regions are cut out to achieve accurate and rapid detection of defect regions in flaw detection film images.

[0021] The present invention constrains the encoder by constructing a teacher model branch that is basically the same as the encoder in structure. During training, the parameters of the teacher model branch are not updated through gradient backpropagation, but are iteratively updated based on the parameters of the encoder. The mean squared error loss is calculated based on the output features of the teacher model branch and the encoder, and the sum of the mean squared error loss and the exponential weighted Focal Loss is used as the total loss. The gradient descent method is used to perform gradient backpropagation and update the parameters of the Transformer distillation model based on the total loss. The construction and update method of the teacher model branch of the present invention can significantly reduce the video memory space occupied during model training, speed up the training speed, save the training cost, and moreover, the calculation method of the total loss adopted in this application can reduce the model training difficulty and improve the adaptability of the model to complex data.

[0022] The method of the present invention can be applied to assist film reviewers in industrial production to quickly detect defect regions for classification and evaluation. By analyzing digital scanned X-ray flaw detection images, rapid discrimination of defective regions is achieved, improving the detection efficiency and saving production time costs.

[0023] The following provides a detailed description of the present invention in conjunction with the accompanying drawings and specific embodiments. Description of the Drawings

[0024] Figure 1 is the training and application flow chart of the present invention.

[0025] Figure 2 is the network structure diagram of the present invention. Specific Embodiments

[0026] As Figures 1 to 2 shown, the method for defect detection of X-ray flaw detection images of nuclear power pipelines based on the Transformer distillation model includes the following steps:

[0027] Step 1: Preprocess the acquired scanned X-ray flaw detection images of nuclear power pipelines, construct a training set and a validation set, and obtain their corresponding mask label sets based on the training set and the validation set. The preprocessing includes numerical mapping, median filtering, and manual marking, and its specific process includes the following steps:

[0028] S1. First, perform numerical mapping on the X-ray flaw detection image I 1 Perform numerical mapping to map the 16-bit depth X-ray flaw detection image obtained by the industrial scanner into an 8-bit depth X-ray flaw detection image, map the pixel values from the original range of 0 to 65536 to the range of 0 to 255, and obtain the mapped X-ray flaw detection image I 2 :

[0029]

[0030] where I 1 is the original 16-bit depth X-ray flaw detection image, I 2 is the converted 8-bit depth X-ray flaw detection image, min(·) is the operation of taking the minimum value of the image, and max(·) is the operation of taking the maximum value of the image.

[0031] S2. Filter the image I 2 after numerical mapping by using the median filtering method. Median filtering is a non-linear filtering processing technology commonly used to remove impulse noise. The standard median filtering algorithm uses a moving window, preferably 3×3, 5×5, or 7×7, etc., moves along the detected image, arranges the gray values of each point in the moving window in an increasing or decreasing manner, and finally replaces the pixel gray value at the center position of the window with the middle value of this sequence to correct the pixel value at the noise point, thereby achieving noise filtering.

[0032] For a two-dimensional image select a filtering window with a size of (2k + 1)×(2k + 1), where k is usually an odd number between 3 and 11. The more image noise points, the larger the value of k. The standard median filter can be defined as:

[0033]

[0034] where A is all the coordinates of the pixel in the filtering window, I 3 is the X-ray flaw detection image after denoising, and Median(·) represents taking the median of all pixels in the filtering window.

[0035] S3. Have professional film interpreters perform on the preprocessed X-ray flaw detection image I 3Manually label the defective areas, generate the corresponding mask image Y based on the results of the manual labeling, and randomly allocate all X-ray flaw detection images and the corresponding mask images to the training set X train and the validation set X eval , corresponding to the set of mask labels Y train and Y eval . Preferably, the random allocation is carried out in a ratio of 8:2.

[0036] Step 2: Construct a Transformer distillation model to segment the defective areas of X-ray flaw detection images. The Transformer distillation model includes an encoder and a decoder. The encoder includes an image patch feature extraction module and a sequence compression multi-head self-attention module. The primary features of the X-ray flaw detection image are extracted through the image patch feature extraction module, and the primary features of the X-ray flaw detection image are enriched through the sequence compression multi-head self-attention module to capture deeper and more complex features to obtain the output features; the decoder is used to reconstruct the image information to obtain the prediction result of the defective area of the X-ray flaw detection image.

[0037] Specifically, the encoder is composed of one image patch feature extraction module and multiple stacked sequence compression multi-head self-attention modules. According to the number and number of heads of the sequence compression multi-head self-attention modules, the encoder can be divided into four types: extremely small, small, basic, and large. Among them, the extremely small type is stacked by 12 sequence compression multi-head self-attention modules, the number of heads of each self-attention module is 3, and the feature dimension is 192; the small type is stacked by 12 sequence compression multi-head self-attention modules, the number of heads of each self-attention module is 6, and the feature dimension is 384; the basic type is stacked by 12 sequence compression multi-head self-attention modules, the number of heads of each self-attention module is 12, and the feature dimension is 768; the large type is stacked by 24 sequence compression multi-head self-attention modules, the number of heads of each self-attention module is 16, and the feature dimension is 1024. In practical applications, the encoder that matches can be flexibly selected according to the number of defect types, the number of labeled images, and the requirements of detection speed.

[0038] The image patch feature extraction module includes three convolutional layers, two batch normalization layers, and two ReLU activation layers. The ReLU activation layer follows the batch normalization layer, and the batch normalization layer and the ReLU activation layer are located between adjacent convolutional layers. The convolutional layer is a conventional convolutional layer with a convolutional kernel size of 15, a stride of 1, and a padding number of 7. Preferably, the output feature of the third convolutional layer is added to the output feature of the first ReLU activation layer to form a set of skip connections for alleviating the problem of gradient disappearance during training. The preprocessed X-ray flaw detection image I 3 becomes a feature map after being processed Among them, h, w, and D respectively represent the height, width, and feature dimension of the feature map. The height and width of the feature map X0 are both 1 / 8 of the original image. To meet the input requirements of the subsequent module, the feature map X0 is flattened in the height and width directions. Specifically, an existing flattening method can be used for flattening. For example, the flatten method is used for flattening. After flattening, the feature map becomes the primary feature sequence of the encoder. Among them, N = h × w, representing the length of the sequence. The image patch feature extraction module completes the extraction of the primary features of the X-ray flaw detection image.

[0039] The sequence compression multi-head self-attention module includes a sequence compression multi-head self-attention layer, a first normalization layer, a multi-layer perceptron, and a second normalization layer connected in sequence. Preferably, the output features of the first normalization layer are added to the input features of the sequence compression multi-head self-attention layer, and the output features of the second normalization layer are added to the output features of the first normalization layer to form two sets of skip connections for alleviating the vanishing gradient problem during training.

[0040] The sequence compression multi-head self-attention layer first uses three linear layers to map the input feature sequence to a query sequence a key sequence and a value sequence where the value of the feature dimension D1 can be determined based on the number of channels of the actual processed image and the parameters actually set by the image patch feature extraction module. Usually, D1 is less than the feature dimension D to save computational space. The key sequence and the value sequence each pass through a convolutional layer for sequence compression to obtain the compressed key sequence and the value sequence Then, multi-head self-attention calculation is performed, that is, the feature dimension is divided into multiple groups of features. Specifically, the feature dimension D is divided into multiple groups of features D1, and each group of features is one head. The query sequence in each head calculates the attention weight with the compressed key sequence, and after scaling by the feature dimension, the compressed value sequence is weighted. After the weighted results of multiple heads are combined, the output of the sequence compression multi-head self-attention layer is obtained. The process can be expressed as:

[0041]

[0042] Among them, is the output of the sequence compression multi-head self-attention layer in the i-th sequence compression multi-head self-attention module; subsequently the feature is normalized through the first normalization layer to obtain The multi-layer perceptron includes two linear layers. The first linear layer first maps the input feature sequence The feature dimension is expanded from D1 to 4D1, and the second linear layer then reduces its feature dimension back to D. After two sets of skip connections in the sequence compression multi-head self-attention module, the output feature sequence of the i-th sequence compression multi-head self-attention module is obtained. The processing process of the multi-layer perceptron can be expressed by the formula:

[0043]

[0044] Finally Then, feature normalization is performed through the second normalization layer to obtain the output features of the current i-th sequence compression multi-head self-attention module. The sequence compression multi-head self-attention module enriches the primary features of the X-ray flaw detection image, enabling the model to capture deeper and more complex features, while avoiding a significant increase in the model's computational complexity and ensuring computational efficiency.

[0045] The decoder includes a decoder module. Preferably, the decoder includes 3 groups of decoder modules. The decoder module includes an upsampling layer and a linear layer. The feature image is enlarged in size through the upsampling layer, and the feature image is dimension-reduced through the linear layer to obtain the final output features with a feature dimension of 2. The final output features are activated by SoftMax to obtain the probability that the current pixel position belongs to the normal or defective area. Specifically, the decoder first converts the 1D feature sequence output by the encoder back into a 2D feature map Specifically, an existing method can be used for restoration, such as using the permute method for restoration. Then, the three groups of decoder modules gradually map the restored features into the segmentation results. The upsampling layer in each group of decoder modules is responsible for upsampling the feature map to twice the current size, and the linear layer is responsible for dimension-reducing the feature dimension. Each group of decoder modules can be expressed by the formula:

[0046]

[0047] where is the output feature from the previous group of decoder modules, Upsample(·) is the upsampling operation for enlarging the feature map size, Linear(·) represents the linear layer, is the output feature of the current decoder module.

[0048] To make full use of the original local features of the image, the feature dimension after dimension reduction by the linear layer of the decoder module is made to correspond one-to-one with the output dimension after the three convolutional layers of the image patch feature extraction module, and the feature map of the image patch feature extraction module is concatenated with the feature map output by the current linear layer to form a long-distance skip connection. At this time, each group of decoder modules can be expressed by the formula:

[0049]

[0050] Among them, is the output feature from the previous group of decoder modules, and Upsample(·) is an upsampling operation used to enlarge the size of the feature map. is the output feature of the j-th convolutional layer in the image patch feature extraction module. Concatenate(·) represents the concatenation operation in the channel dimension, and Linear(·) represents the linear layer. is the output feature of the current decoder module.

[0051] The output feature dimension of the last linear layer of the decoder is 2. After being activated by SoftMax, the values therein represent the probabilities that the current pixel position belongs to the normal or defective area. The probability output by the decoder here is denoted as The decoder restores the detailed information layer by layer and remains sensitive to features of various scales, and finally obtains a high-precision segmentation result.

[0052] Step 3: Construct a teacher model branch. The teacher model branch is basically the same as the encoder in structure, except that the sequence compression multi-head self-attention module of the encoder is replaced by a conventional multi-head self-attention module. It is also stacked by an image patch feature extraction module and a multi-head self-attention module, and also receives the preprocessed X-ray flaw detection image as the input and finally outputs the extracted X-ray flaw detection image feature X t . Its multi-head self-attention module includes a multi-head self-attention layer, a first normalization layer, a multi-layer perceptron, and a second normalization layer connected in sequence. Compared with the sequence compression multi-head self-attention module, the series compression operation is removed, and other calculation processes are the same.

[0053] Step 4: Train the Transformer distillation model and the teacher model branch based on the training set, the validation set, and their corresponding mask label sets. During the training, the parameters of the teacher model branch are not updated through gradient backpropagation, but are iteratively updated based on the parameters of the encoder. Specifically, the iterative calculation process of the teacher model branch is as follows:

[0054] Θ′ t = 0.5(Θ t + Θ s ),

[0055] where, Θ t ′ represents the updated parameters of the teacher model branch, Θ t represents the parameters of the teacher model branch before update, and Θ s represents the parameters of the current encoder.

[0056] During training, the output features of the teacher model branch will be used to supervise the feature output of the encoder. The supervision method is to use the output features of the teacher model branch as the ground truth, and calculate the mean squared error loss with the output features of the encoder as the predicted values:

[0057]

[0058] where X s and X t represent the output features of the encoder and the teacher model branch respectively.

[0059] In addition to the mean squared error loss , an exponential weighted Focal Loss is also added. The exponential weighted Focal Loss is based on the cross-entropy loss. Considering that the proportion of the defect area in the whole X-ray flaw detection image is small, that is, there is a significant class imbalance problem in the dataset. Therefore, a weight coefficient is added in front of the cross-entropy loss to increase the weight of the class with fewer samples. In addition, considering that the easy-to-separate samples (i.e., samples with high confidence) have very little improvement effect on the model. Therefore, in order to make the model pay more attention to the difficult-to-separate samples, the weight of the high-confidence samples is further reduced by introducing an exponential coefficient. The calculation process of the exponential weighted Focal Loss is:

[0060]

[0061] where is the probability that the prediction at a single pixel location is a normal area; y comes from the set of mask labels in the training set, that is, y ∈ Y train , that is, the ground truth label of this pixel location obtained through manual annotation. Among them, the normal area is 0, and the defect area is 1; α is the weight coefficient used to alleviate the class imbalance problem. Preferably, α can be set to 4; γ is the exponential coefficient that makes the model pay more attention to difficult samples. Preferably, γ can be set to 2.

[0062] During training, the total loss finally used is the sum of the mean squared error loss and the exponential weighted Focal Loss . During training, the gradient descent method is used to perform gradient backpropagation and model parameter update based on the calculated total loss. For example, the SGD method is used for gradient backpropagation and model parameter update.

[0063] During training, the training set (X train , Y train ) is repeatedly used to train the model. All samples in the training set are passed through the model once and recorded as one round. After each round of training, the X eval of the validation set is input into the model, and the model output probability is compared with the Y eval of the validation set.Calculate the total loss of the validation set, and stop training when the total loss of the validation set no longer decreases for three consecutive rounds.

[0064] Step 5: After training the Transformer distillation model, use the trained model to detect the preprocessed X-ray flaw detection images to be detected. First, use the same preprocessing method as for processing the training data to perform numerical mapping and median filtering on the X-ray flaw detection images to be detected, and then input the images into the trained Transformer distillation model network to obtain the probability map of the defect area. Among them, the teacher model branch does not participate in the calculation during testing, and the output features of the encoder are only used for the decoder to generate the segmentation result. The probability map of the defect area output by the decoder contains two channels. For each pixel position, if the probability value of the first channel is larger, the category of this pixel position is the normal area; if the probability value of the second channel is larger, the category of this pixel position is the defect area, and thus a binary map indicating the defect area is obtained. This binary map can be used to assist the film reviewers in nuclear power pipeline flaw detection.

[0065] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of these features. In the description of the present invention, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0066] In the present invention, the terms "an embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0067] Although the above embodiments have been shown and described, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Any changes, modifications, substitutions, and variations made by those of ordinary skill in the art to the above embodiments are within the protection scope of the present invention.

Claims

1. A method for detecting defects in X-ray flaw detection images of nuclear power pipelines based on a Transformer distillation model, characterized in that, It includes the following steps: Step 1: Preprocess the acquired scanned X-ray flaw detection images of nuclear power pipelines, construct a training set and a validation set, and obtain their corresponding mask label sets based on the training set and the validation set; Step 2: Construct a Transformer distillation model to segment the defect areas of X-ray flaw detection images. The Transformer distillation model includes an encoder and a decoder. The encoder includes an image patch feature extraction module and a sequence compression multi-head self-attention module. The primary features of the X-ray flaw detection images are extracted through the image patch feature extraction module, and the primary features of the X-ray flaw detection images are feature enhanced through the sequence compression multi-head self-attention module to obtain output features; The decoder is used to obtain the prediction results of the defect areas of the X-ray flaw detection images based on the output features extracted by the encoder; Step 3: Construct a teacher model branch, which includes an image patch feature extraction module and a multi-head self-attention module. The primary features are extracted through the image patch feature extraction module, and the output features are obtained through the multi-head self-attention module; Step 4: Train the Transformer distillation model and the teacher model branch based on the training set, the validation set and their corresponding mask label sets. During training, the parameters of the teacher model branch are not updated through gradient backpropagation, but are iteratively updated based on the parameters of the encoder. Calculate the mean squared error loss based on the output features of the teacher model branch and the encoder, and use the sum of the mean squared error loss and the exponential weighted Focal Loss as the total loss. Use the gradient descent method to perform gradient backpropagation and update the parameters of the Transformer distillation model based on the total loss; Step 5: After the Transformer distillation model is trained, use the trained model to detect the preprocessed X-ray flaw detection images to be detected.

2. The method for detecting defects in X-ray flaw detection images of nuclear power pipelines based on the Transformer distillation model according to claim 1, characterized in that: The image patch feature extraction module includes three convolutional layers, two batch normalization layers and two ReLU activation layers. The ReLU activation layer follows the batch normalization layer, and the batch normalization layer and the ReLU activation layer are located between two adjacent convolutional layers.

3. The method for detecting defects in X-ray flaw detection images of nuclear power pipelines based on the Transformer distillation model according to claim 2, wherein: The output features of the third convolutional layer are added to the output features of the first ReLU activation layer to form a set of skip connections.

4. The method for detecting defects in X-ray flaw detection images of nuclear power pipelines based on the Transformer distillation model according to claim 1, wherein: The sequence compression multi-head self-attention module includes a sequence compression multi-head self-attention layer, a first normalization layer, a multi-layer perceptron, and a second normalization layer connected in sequence.

5. The method for defect detection of X-ray flaw detection images of nuclear power pipelines based on the Transformer distillation model according to claim 4, characterized in that: The sequence compression multi-head self-attention layer first uses three linear layers to map the output features of the image patch feature extraction module into a query sequence, a key sequence and a value sequence. The key sequence and the value sequence are each compressed through a convolutional layer, and then multi-head self-attention calculation is performed based on the query sequence and the compressed key sequence and value sequence.

6. The method for defect detection of X-ray flaw detection images of nuclear power pipelines based on the Transformer distillation model according to claim 4, characterized in that: The output features of the first normalization layer are added to the input features of the sequence compression multi-head self-attention layer, and the output features of the second normalization layer are added to the output features of the first normalization layer to form two sets of skip connections.

7. The method for defect detection of X-ray flaw detection images of nuclear power pipelines based on the Transformer distillation model according to claim 2, characterized in that: The decoder includes a decoder module, and the decoder module includes an upsampling layer and a linear layer. The feature image is enlarged in size through the upsampling layer, and the feature image is dimensionally reduced through the linear layer to obtain a final output feature with a feature dimension of 2. The final output feature is activated by SoftMax to obtain the probability that the current pixel position belongs to the normal or defective area.

8. The method for defect detection of X-ray flaw detection images of nuclear power pipelines based on the Transformer distillation model according to claim 7, wherein: The feature dimension after dimensional reduction by the linear layer of the decoder module corresponds one-to-one with the output dimension after passing through the three convolutional layers of the image patch feature extraction module, and the feature map of the image patch feature extraction module is concatenated with the feature map output by the current linear layer to form a long-range skip connection.

9. The method for defect detection of X-ray flaw detection images of nuclear power pipelines based on the Transformer distillation model according to claim 1, characterized in that: The iterative calculation process of the teacher model branch is Θ′ t = 0.5(Θ t + Θ s ), where Θ t ′ represents the parameters of the updated teacher model branch, Θ t represents the parameters of the teacher model branch before update, and Θ s represents the parameters of the current encoder.

10. The method for defect detection of X-ray flaw detection images of nuclear power pipelines based on the Transformer distillation model according to claim 1, characterized in that: The calculation process of the exponential weighted Focal Loss is as follows where is the probability that the prediction at a single pixel location is a normal region, y comes from the set of mask labels in the training set, i.e., y ∈ Y train , represents the ground truth label at this pixel location, α is the weight coefficient used to alleviate the class imbalance problem, and γ is the exponential coefficient that makes the model pay more attention to hard samples.

Citation Information

Patent Citations

  • Nuclear power pipeline defect detection method based on multi-scale pyramid structure

    CN111899225A

  • Transformer substation equipment defect image detection method based on changeable patch

    CN115937091A

  • Railway wagon steel floor damage detection method based on knowledge distillation

    CN116703819A

  • Aircraft skin surface defect detection method based on lightweight neural network

    CN117474914A