A Two-Stage Intelligent Recognition Method and System for X-Ray Images of Weld Defects with Unequal Thickness Based on Reverse Learning
Through the two-stage intelligent identification method of reverse learning, weld defect identification network is trained using positive and negative samples, which solves the problem of dependence on a large amount of training data in the existing technology, and achieves efficient and accurate identification of weld defects.
Patent Information
- Application Number
- CN202410658124.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-05-24
AI Technical Summary
Existing methods for automatic identification of weld defect X-ray images based on deep learning rely heavily on a large amount of training data, especially when obtaining defective weld X-ray image samples, resulting in low recognition accuracy and efficiency.
The two-stage intelligent recognition method of reverse learning is adopted. In the first stage, the weld defect potential area identification network is trained simultaneously by positive and negative samples to identify the defect potential area; in the second stage, only negative samples are used to train the accurate identification network, and slice images are recognized to achieve accurate identification of defects.
This method can realize intelligent, efficient and accurate identification of weld defects with a small number of positive samples, avoiding the dependence of deep learning models on a large number of positive samples, and improving the recognition accuracy and efficiency.
Smart Images

Figure CN118587174B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of nondestructive testing, and more specifically, to a two-stage intelligent recognition method and system for X-ray images of unequal-thickness weld defects with reverse learning. Background Art
[0002] Nondestructive testing technology is widely used in the detection of weld defects. Among them, X-ray digital imaging technology is mainly used for imaging the welds of various aluminum and steel materials due to its high detection accuracy. Then, experienced experts judge the formed weld images to determine whether there are defects, or automatic judgment is carried out on the formed weld images through artificial intelligence algorithms. The existing automatic recognition methods for X-ray images of weld defects based on artificial intelligence rely heavily on training data sets. For some fields, it is very difficult to obtain a large number of defective weld X-ray images. Therefore, the existing methods based on deep learning of big data are not applicable, and the accuracy of the existing meta-learning and transfer learning methods that rely on a small number of samples still needs to be improved and cannot achieve the application purpose.
[0003] In order to be able to use a small number of defective weld X-ray image samples to achieve intelligent and accurate recognition of weld defects, a method different from conventional deep learning needs to be studied. Conventional deep learning models realize the recognition of test images by directly learning the feature patterns of positive samples. In order to learn the accurate feature patterns of positive samples, a large number of positive samples are required as training data. The reverse learning method can indirectly judge the existence of the target by learning the feature patterns of negative samples. In addition, since only a part of the weld images obtained by X-ray digital imaging technology is the weld area, and the other parts are useless background areas, those useless background areas not only generate potential misrecognition areas for target recognition, but also reduce the recognition efficiency. Therefore, the present invention proposes a method that can intelligently and accurately recognize the X-ray images of weld defects by only using a small number of positive samples plus some negative samples. Summary of the Invention
[0004] In view of this, the present invention provides a two-stage intelligent recognition method and system for X-ray images of unequal-thickness weld defects with reverse learning, achieving the purpose of intelligent, efficient, and accurate recognition of weld defects.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions:
[0006] A two-stage intelligent recognition method for X-ray images of unequal-thickness weld defects with reverse learning, comprising:
[0007] Obtain weld images and establish a first data set; the images with defects in the weld images are positive samples, and the images without defects in the weld images are negative samples; establish a second data set according to the positive samples;
[0008] Input the first data set into a preset weld defect potential area recognition model to train the weld defect potential area recognition model; input the second data set into a preset weld defect precise recognition model to train the weld defect precise recognition model.
[0009] Obtain a weld image to be recognized, input the weld image to be recognized into the trained weld defect potential area recognition model to identify the defect potential area of the weld image to be recognized; slice the weld defect potential area to obtain a sliced image, and use the sliced image as the input of the trained weld defect precise recognition model to identify and mark the sliced image with a defect potential area.
[0010] Preferably, it further includes marking the defect potential area of the weld for the positive sample; performing a slicing operation on the negative sample.
[0011] Preferably, the preset weld defect potential area recognition model is specifically a YOLOv8 object detection network model, and the parameters trained on the ImageNet data set are used as the pre-training parameters of the YOLOv8 object detection network model.
[0012] Preferably, the weld defect precise recognition model specifically includes: a generation module, a re-encoding module, and a discrimination module. The generation module obtains the sliced feature map of the sliced image and reconstructs it to obtain a sliced reconstruction map; the re-encoding module extracts the reconstructed feature map of the sliced reconstruction map; the discrimination module inputs the reconstructed feature map and the sliced image to judge the reconstructed feature map and the sliced image, and helps the generation module learn the features of normal weld images.
[0013] Preferably, the generation module specifically includes an encoding sub-module and a decoding sub-module, and the specific structure is as follows:
[0014] The encoding sub-module includes a first layer as the input layer with an input channel number of 3; a second layer as a convolutional layer with an input channel number of 3, an output channel number of 32, a convolutional kernel size of 3*3, and a stride of 2*2; a third layer as a Relu activation layer; a fourth layer as a convolutional layer with an input channel number of 32, an output channel number of 64, a convolutional kernel size of 3*3, and a stride of 2*2; a fifth layer as a batch normalization layer; a sixth layer as a Relu activation layer; a seventh layer as a convolutional layer with an input channel number of 64, an output channel number of 128, a convolutional kernel size of 3*3, and a stride of 2*2;
[0015] The sliced feature map f output by the encoding sub-module is expressed as follows:
[0016] f = G E (x; θe )
[0017] Among them, θ e represents the network parameters corresponding to the encoding sub-module, and G E (x; θ e ) represents the encoding sub-module, and x represents the input training sample;
[0018] The decoding sub-module includes a first layer of transposed convolution layer with an input channel number of 128, an output channel number of 64, a convolution kernel size of 3*3, and a stride of 2*2; the second layer is a batch normalization layer; the third layer is a Relu activation layer; the fourth layer is a transposed convolution layer with an input channel number of 64, an output channel number of 32, a convolution kernel size of 3*3, and a stride of 2*2; the fifth layer is a batch normalization layer; the sixth layer is a Relu activation layer; the seventh layer is a transposed convolution layer with an input channel number of 32, an output channel number of 3, a convolution kernel size of 3*3, and a stride of 2*2;
[0019] The reconstructed defect potential area slice map output by the decoding sub-module is expressed as follows:
[0020]
[0021] Among them, θ d represents the network parameters corresponding to the decoding sub-module, and G D (f; θ d ) represents the decoding sub-module.
[0022] Preferably, the re-encoding module specifically includes:
[0023]
[0024] Among them, θ e1 represents the network parameters corresponding to the re-encoding module, represents the re-encoding module, represents the reconstructed feature map output by the re-encoding module.
[0025] Preferably, the discriminant module specifically includes a scale-one discriminant module and a scale-two discriminant module, and the specific structure is as follows:
[0026] The scale-one discrimination module includes: The first layer is the input layer with 3 input channels; the second layer is the convolutional layer with 3 input channels, 32 output channels, a convolutional kernel size of 3*3, and a stride of 2*2; the third layer is the Relu activation layer; the fourth layer is the convolutional layer with 32 input channels, 64 output channels, a convolutional kernel size of 3*3, and a stride of 2*2; the fifth layer is the batch normalization layer; the sixth layer is the Relu activation layer; the seventh layer is the convolutional layer with 64 input channels, 128 output channels, a convolutional kernel size of 3*3, and a stride of 2*2; the eighth layer is the batch normalization layer; the ninth layer is the Relu activation layer; the tenth layer is the convolutional layer with 128 input channels, 1 output channel, a convolutional kernel size of 3*3, and a stride of 1*1; the eleventh layer is the Sigmoid activation layer; The output of the scale-one discrimination module is as follows:
[0027]
[0028] where l 1 results in 0 or 1. When it is 0, it represents the reconstructed feature map, and when it is 1, it represents the sliced image. θ D1 represents the network parameters corresponding to the scale-one discrimination module, is the scale-one discrimination module;
[0029] The scale-two discrimination module includes: The first layer is the input layer with 64 channels; the second layer is the convolutional layer with 64 input channels, 128 output channels, a convolutional kernel size of 3*3, and a stride of 2*2; the third layer is the batch normalization layer; the fourth layer is the Relu activation layer; the fifth layer is the convolutional layer with 128 input channels, 1 output channel, a convolutional kernel size of 3*3, and a stride of 1*1; the sixth layer is the Sigmoid activation layer; The output of the scale-two discrimination module is as follows:
[0030] l 2 = D 2 (G E-2 (x), G D-2 (f); θ D2 );
[0031] where l 2 results in 0 or 1. When it is 0, it represents the reconstructed feature map, and when it is 1, it represents the sliced image. θ D2 represents the network parameters corresponding to the scale-two discrimination module, G E-2 represents the network module of the 4th layer and the layers before the 4th layer in the encoding sub-module of the generation module, G D-2 represents the network module of the 4th layer and the layers before it in the decoding sub-module, D 2 (G E-2 (x), G D-2 (f); θ D2) is the scale-two discrimination module.
[0032] A two-stage intelligent recognition system for X-ray images of unequal-thickness weld defects with reverse learning, comprising:
[0033] A dataset establishment module, which acquires weld images and establishes a first dataset; the images with defects in the weld images are positive samples, and the images without defects in the weld images are negative samples; a second dataset is established according to the positive samples;
[0034] A model training module, which inputs the first dataset into a preset weld defect potential area recognition model to train the weld defect potential area recognition model; inputs the second dataset into a preset weld defect precise recognition model to train the weld defect precise recognition model;
[0035] A region recognition module, which acquires a weld image to be recognized, inputs the weld image to be recognized into the trained weld defect potential area recognition model to recognize the defect potential area of the weld image to be recognized; slices the weld defect potential area to obtain slice images, and takes the slice images as the input of the trained weld defect precise recognition model to recognize and mark the slice images with defect potential areas therein.
[0036] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a two-stage intelligent recognition method and system for X-ray images of unequal-thickness weld defects with reverse learning. Based on deep learning technology, aiming at the problems existing in the application of existing deep learning-based methods, a new automatic recognition method for weld defects is proposed. In the first stage, both positive samples and negative samples are used to train the weld defect potential area recognition network to recognize the defect potential areas in the original weld images. In the second stage, only negative samples are used to train the defect recognition network, and precise defect recognition is carried out in the second stage on the basis of the defect potential areas recognized in the first stage. This not only avoids the disadvantage of relying on a large number of positive samples during the training of the deep learning model, but also can greatly improve the recognition accuracy and recognition efficiency of weld defects. Description of the Drawings
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0038] Figure 1 It is a schematic diagram of an intelligent recognition method for X-ray images of unequal-thickness weld defects with reverse learning provided by the present invention;
[0039] Figure 2 This is an example of the first dataset sample provided by the present invention;
[0040] Figure 3 This is an example of the second dataset sample provided by the present invention;
[0041] Figure 4 This is a schematic diagram of the defect potential area recognition model in the first stage provided by the present invention;
[0042] Figure 5 This is a schematic diagram of an image slice in the weld defect precise recognition model in the second stage provided by the present invention;
[0043] Figure 6 This is a loss function graph of the weld defect precise recognition model in the second stage provided by the present invention;
[0044] Figure 7 This is a general schematic diagram of an intelligent recognition model for X-ray images of unequal-thickness weld defects with reverse learning provided by the present invention. Detailed implementation manners
[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0046] The embodiment of the present invention discloses a two-stage intelligent recognition method for X-ray images of unequal-thickness weld defects with reverse learning, as Figure 1 shown, including:
[0047] Obtain a weld image and establish a first dataset; the image with defects in the weld image is a positive sample, and the image without defects in the weld image is a negative sample; establish a second dataset according to the positive samples;
[0048] Input the first dataset into a preset weld defect potential area recognition model to train the weld defect potential area recognition model; input the second dataset into a preset weld defect precise recognition model to train the weld defect precise recognition model;
[0049] Obtain a weld image to be recognized, input the weld image to be recognized into the trained weld defect potential area recognition model to recognize the defect potential area of the weld image to be recognized; slice the defect potential area of the weld to obtain a sliced image, and use the sliced image as the input of the trained weld defect precise recognition model to recognize and mark the sliced image with a defect potential area therein.
[0050] The first data set is unequal-thickness weld images obtained by an X-ray digital imaging device. Some of these images contain defects, while others do not, but the normal weld images without defects are dominant. Sample examples are as follows Figure 2 shown; The second data set is made based on the normal weld image samples in the first data set. The potential defect areas in the normal weld images are sliced to obtain sliced tile samples of the potential defect areas of the welds. Sample examples are as follows Figure 3 shown.
[0051] In a specific embodiment, it also includes marking the potential defect areas of the welds for positive samples; performing slicing operations on negative samples. Slicing the potential defect areas of the welds can effectively improve the recognition accuracy of small defects, as follows Figure 5 shown. The potential defect areas of the welds are sliced into image blocks of a fixed size. Through preliminary experiments, the results show that when the size of each slice is 64×64 pixel points, the effect is better. Compared with non-slicing, slicing can effectively retain the detailed features of small defects. As shown in the following figure, when not sliced, since the size of the potential defect area of the weld is relatively large, it needs to be resized first after being input into the defect recognition model to meet the requirements of the conventional recognition model for the pixel size of the input sample, which is 640×640. After slicing, since the potential defect area of the weld is sliced into multiple small tiles of 64×64, therefore, it can be input into the recognition model for defect recognition without performing the resize operation.
[0052] In a specific embodiment, the preset potential weld defect area recognition model is specifically the YOLOv8 object detection network model, and the parameters trained on the ImageNet data set are used as the pre-training parameters of the YOLOv8 object detection network model. Since the targets in the potential defect areas of the welds are generally relatively large, the existing YOLOv8 object detection model is directly used as the potential weld defect area recognition model of the present invention, as follows Figure 4 shown. The YOLOv8 object detection model can not only ensure accuracy but also has very good real-time performance, which is convenient for deployment on embedded devices.
[0053] In a specific embodiment, in order to eliminate the problem of poor model generalization performance caused by insufficient positive samples, the present invention proposes a weld defect recognition model that combines reverse learning with slicing technology. The potential weld defect regions detected by the first-stage potential weld defect region recognition model are used as the input of the weld defect recognition model in this stage. Then, the input potential weld defect regions are sliced, and the trained model is used to recognize each slice. After comprehensively evaluating the recognition results of each slice, it is finally determined whether there are defects in the potential weld regions of the defect, and the defect positions are displayed. Aiming at the problem that the conventional weld defect recognition model overly relies on the number of positive samples, since normal defect-free samples can be easily obtained, it is entirely possible to bypass the limitation of insufficient positive sample quantity. This model is based on a generative adversarial network and has been improved on this basis. The weld defect precise recognition model specifically includes: a generation module, a re-encoding module, and a discrimination module. The generation module obtains the slice feature map of the slice image and reconstructs it to obtain a slice reconstruction map; the re-encoding module extracts the reconstruction feature map of the slice reconstruction map; the discrimination module inputs the reconstruction feature map and the slice image to judge the reconstruction feature map and the slice image, and helps the generation module learn the features of normal weld images.
[0054] In a specific embodiment, the generation module specifically includes an encoding sub-module and a decoding sub-module. The encoding sub-module is used to obtain the slice feature map representation of the input slice image block, which consists of 3 convolutional layers and their corresponding activation functions and batch normalization. The decoding sub-module is used to reconstruct the image, reconstruct the slice feature map obtained by the encoding sub-module, and generate a slice reconstruction map that is realistic to the slice image block of the original input potential defect region. This sub-module consists of 3 transposed convolutional layers and corresponding batch normalization and activation functions. The specific structure is as follows:
[0055] The encoding sub-module includes: the first layer is the input layer with an input channel number of 3; the second layer is the convolutional layer (the first convolutional layer) with an input channel number of 3, an output channel number of 32, a convolutional kernel size of 3*3, and a stride of 2*2; the third layer is the Relu activation layer; the fourth layer is the convolutional layer (the second convolutional layer) with an input channel number of 32, an output channel number of 64, a convolutional kernel size of 3*3, and a stride of 2*2; the fifth layer is the batch normalization layer; the sixth layer is the Relu activation layer; the seventh layer is the convolutional layer (the third convolutional layer) with an input channel number of 64, an output channel number of 128, a convolutional kernel size of 3*3, and a stride of 2*2;
[0056] The slice feature map f output by the encoding sub-module is expressed as follows:
[0057] f = G E (x; θ e );
[0058] where, θ eRepresents the network parameters corresponding to the encoding sub-module, G E (x; θ e ) represents the encoding sub-module, where x represents the input training sample;
[0059] The decoding sub-module includes a first layer that is a transposed convolution layer (the first transposed convolution layer), with an input channel number of 128, an output channel number of 64, a convolution kernel size of 3*3, and a stride of 2*2; the second layer is a batch normalization layer; the third layer is a Relu activation layer; the fourth layer is a transposed convolution layer (the second transposed convolution layer), with an input channel number of 64, an output channel number of 32, a convolution kernel size of 3*3, and a stride of 2*2; the fifth layer is a batch normalization layer; the sixth layer is a Relu activation layer; the seventh layer is a transposed convolution layer (the third transposed convolution layer), with an input channel number of 32, an output channel number of 3, a convolution kernel size of 3*3, and a stride of 2*2;
[0060] The reconstructed map of the defect potential region slices output by the decoding sub-module is expressed as follows:
[0061]
[0062] where θ d represents the network parameters corresponding to the decoding sub-module, G D (f; θ d ) represents the decoding sub-module.
[0063] In a specific embodiment, the re-encoding module has the same structure as the encoding sub-module in the generation module and is used to obtain an accurate feature map representation of the reconstructed slice image block, specifically including:
[0064]
[0065] where θ e1 represents the network parameters corresponding to the re-encoding module, represents the re-encoding module, represents the reconstructed feature map output by the re-encoding module.
[0066] In a specific embodiment, the discrimination module can enable the generation module to only learn the features of normal weld images. This module includes two scales, corresponding to the scale-one discrimination module and the scale-two discrimination module respectively. Each scale module is composed of a convolution layer and its corresponding activation function and batch normalization. The specific structure is as follows:
[0067] The scale-one discrimination module includes: The first layer is the input layer with 3 input channels; the second layer is the convolutional layer with 3 input channels, 32 output channels, a convolutional kernel size of 3*3, and a stride of 2*2; the third layer is the Relu activation layer; the fourth layer is the convolutional layer with 32 input channels, 64 output channels, a convolutional kernel size of 3*3, and a stride of 2*2; the fifth layer is the batch normalization layer; the sixth layer is the Relu activation layer; the seventh layer is the convolutional layer with 64 input channels, 128 output channels, a convolutional kernel size of 3*3, and a stride of 2*2; the eighth layer is the batch normalization layer; the ninth layer is the Relu activation layer; the tenth layer is the convolutional layer with 128 input channels, 1 output channel, a convolutional kernel size of 3*3, and a stride of 1*1; the eleventh layer is the Sigmoid activation layer; The scale-one discrimination module inputs the training sample x and The network performs layer-by-layer convolution and related operations on it, and finally outputs the corresponding label (real slice image block or reconstructed slice image block). The specific output is as follows:
[0068]
[0069] where, l 1 The result of is 0 or 1. When it is 0, it represents the reconstructed feature map, and when it is 1, it represents the slice image. θ D1 represents the network parameters corresponding to the scale-one discrimination module, is the scale-one discrimination module;
[0070] The scale-two discrimination module includes: The first layer is the input layer with 64 channels; the second layer is the convolutional layer (the first convolutional layer) with 64 input channels, 128 output channels, a convolutional kernel size of 3*3, and a stride of 2*2; the third layer is the batch normalization layer; the fourth layer is the Relu activation layer; the fifth layer is the convolutional layer (the second convolutional layer) with 128 input channels, 1 output channel, a convolutional kernel size of 3*3, and a stride of 1*1; the sixth layer is the Sigmoid activation layer; The scale-two discrimination module performs layer-by-layer convolution and related operations on it, and finally outputs the corresponding label (real slice image block or reconstructed slice image block). The specific output is as follows:
[0071] l 2 = D 2 (G E-2 (x), G D-2 (f); θ D2 );
[0072] where, l 2 The result of is 0 or 1. When it is 0, it represents the reconstructed feature map, and when it is 1, it represents the slice image. θ D2 represents the network parameters corresponding to the scale-two discrimination module, G E-2Denote the network modules before and including the 4th layer of the encoding sub-module in the generation module, G D-2 Denote the network modules before and including the 2nd transposed convolutional layer of the decoding sub-module, D 2 (G E-2 (x), G D-2 (f); θ D2 ) is the scale-two discriminant module.
[0073] A two-stage intelligent recognition system for X-ray images of unequal-thickness weld defects based on reverse learning, comprising:
[0074] A dataset establishment module, which acquires weld images and establishes a first dataset; the images with defects in the weld images are positive samples, and the images without defects in the weld images are negative samples; a second dataset is established based on the positive samples;
[0075] A model training module, which inputs the first dataset into a preset weld defect potential area recognition model to train the weld defect potential area recognition model; inputs the second dataset into a preset weld defect precise recognition model to train the weld defect precise recognition model;
[0076] A region recognition module, which acquires a weld image to be recognized, inputs the weld image to be recognized into the trained weld defect potential area recognition model to recognize the defect potential area of the weld image to be recognized; slices the weld defect potential area to obtain slice images, and uses the slice images as the input of the trained weld defect precise recognition model to recognize and mark the slice images with defect potential areas therein.
[0077] In a specific embodiment, the second dataset constructed previously is used to effectively train the weld defect precise recognition model in the second stage. As Figure 6 shown, the training process includes a total of 4 loss functions, namely adversarial loss, semantic loss, encoding loss, and discriminant loss.
[0078] The adversarial loss represents the L2 distance - Euclidean distance (adversarial loss 1) between the "true and false" feature maps of the input slice image patch and the reconstructed generated image patch. Since the present invention calculates the adversarial loss for two different scales of feature maps simultaneously, the adversarial loss also includes the L2 distance - Euclidean distance (adversarial loss 2) between the "true and false" feature maps after the input slice image patch passes through the second convolutional module of the encoding sub-module and before the second transposed convolutional module of the decoding sub-module of the reconstructed generated image patch. The generation module of the model is trained by minimizing the distance between these two different scales of "true and false" feature maps, so that the image patch reconstructed and generated by the generation module can pass for real. The specific adversarial loss 1 and adversarial loss 2 are respectively expressed as follows:
[0079]
[0080] Among them, f d represents the part of the discrimination network module for scale one and scale two that outputs the "true / false" feature map. For the scale one discrimination network module, f d specifically refers to the fourth layer and the network before it in the scale one discrimination module. For the scale two discrimination network module, f d specifically refers to the first layer of the scale two discrimination module, and p x represents the data distribution of the input sliced image patch. f 2 represents the feature map output by the second convolutional layer of the encoding sub-module, and G D-2 represents the part of the module used to output the feature map of the second transposed convolutional layer in the decoding sub-module. is the expected value, and G E is the encoding sub-module.
[0081] The semantic loss represents the L1 distance - Manhattan distance between the original image and the reconstructed image, and is used to cooperate with the adversarial loss to penalize the generation network module. The specific expression of this semantic loss is as follows:
[0082]
[0083] G D is the decoding sub-module. The encoding loss represents the L2 distance - Euclidean distance between two different-scale feature maps obtained by the encoding sub-module and the re-encoding module in the generation module. These two different-scale feature maps are the feature maps output by the second convolutional layer and the third convolutional layer of the encoding sub-module and the re-encoding module respectively. The encoding losses 1 and 2 corresponding to these two different scales are expressed as follows:
[0084]
[0085] Among them, G E-2 represents the network module including the second convolutional layer and the previous layers in the encoding sub-module of the generation module, and E -2 represents the network module including the second convolutional layer and the previous layers in the re-encoding module.
[0086] The discrimination loss represents the cross-entropy loss between the original image and the reconstructed image, and between a certain feature map of the original image and a certain feature map of the reconstructed image. It is used to judge the "true / false" of the input sliced image patch and the reconstructed generated image patch (corresponding to discrimination loss 1), and at the same time to judge the "true / false" of the input sliced image patch after passing through the second convolutional module of the encoding sub-module and the second transposed convolutional module of the decoding sub-module before the reconstructed generated image patch (corresponding to discrimination loss 2). The discrimination losses 1 and 2 are expressed as follows:
[0087]
[0088]
[0089] Among them, p i and p i2 respectively represent the "true or false" class probabilities output by the sigmoid function of the last layer of the scale-one discriminant module and the scale-two discriminant module. N represents the total number of training samples, and y i is the class label value, which is 0 or 1.
[0090] Adversarial Loss: Adversarial loss is a loss function used in generative adversarial networks (GANs). By training the adversarial process between the generator and the discriminator, the realism and quality of the images generated by the generator can be improved.
[0091] Semantic Loss: Semantic loss is used to ensure that the generated images are semantically consistent with the target images, that is, the generated images can accurately express the content and features of the target images.
[0092] Encoding Loss: Encoding loss is used to keep the encoding information of the generated images consistent with the encoding information of the original images, ensuring that the generated images are similar to the original images in the feature space.
[0093] Discriminative Loss: Discriminative loss is used to train the discriminator network to help the discriminator accurately distinguish between generated images and real images, thereby improving the quality and realism of the generated images.
[0094] Adopting these four loss functions can constrain and optimize the model at different levels, helping the model better learn and generate images that meet the requirements. By comprehensively using these loss functions, the performance and generalization ability of the model can be improved, enabling the model to achieve better results in the defect recognition task.
[0095] In a specific embodiment, it also includes the construction and testing of the overall model: combining the models of the above two stages as follows Figure 7As shown, the X-ray image of the weld to be inspected is used as the input of the first-stage defect potential area recognition model, and the defect potential area output by the first-stage defect potential area recognition model is used as the input of the second-stage defect precise recognition model. Among them, the second-stage defect precise recognition model first slices the input defect potential area, and then performs encoding and re-encoding operations on all sliced images. The L1 distance is measured between the sliced feature map distribution obtained after the sliced image passes through the encoding sub-module and the feature map distribution obtained after the sliced reconstruction map passes through the re-encoding module (the measured distance is between the two feature map distributions). Finally, if all the measured values are not greater than the threshold, it indicates that there is no defect in the defect potential area, thus determining that the X-ray image of the weld to be inspected is normal; if a certain measured value is greater than the threshold, it indicates that there is a defect in the defect potential area, thus determining that there is a defect in the X-ray image of the weld to be inspected. Further, the position of the slice with the defect in the defect potential area is located, and the position of the slice in the original X-ray image of the weld to be inspected is determined by the repositioning method. Thus, the automatic recognition of defects in the entire weld X-ray image is completed.
[0096] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and reference can be made to the description in the method part for related parts.
[0097] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A two-stage intelligent recognition method for X-ray images of unequal thickness weld defects based on reverse learning, characterized in that: include: Acquire a weld image and establish a first data set; The image with defects in the weld image is a positive sample, and the image without defects in the weld image is a negative sample; Establish a second data set based on the negative samples; the first data set includes positive samples and negative samples; Inputting the first data set into a preset weld defect potential area recognition model to train the weld defect potential area recognition model; inputting the second data set into a preset weld defect precise recognition model to train the weld defect precise recognition model; Acquire a weld image to be identified, input the weld image to be identified into the trained weld defect potential area identification model to identify and obtain the defect potential area of the weld image to be identified; Slice the potential weld defect area to obtain a slice image, use the slice image as input of the trained weld defect accurate recognition model, and identify and mark the slice image in which the potential defect area exists; The weld defect accurate identification model specifically includes: a generation module, a recoding module, and a discrimination module. The generation module obtains the slice feature map of the slice image and reconstructs it to obtain a slice reconstruction map; the recoding module extracts the reconstruction feature map of the slice reconstruction map; the discrimination module inputs the reconstruction feature map and the slice image, judges the reconstruction feature map and the slice image, and helps the generation module learn the characteristics of a normal weld image; The generation module specifically includes an encoding submodule and a decoding submodule, and the specific structure is as follows: The encoding submodule includes: the first layer is an input layer, the number of input channels is 3; the second layer is a convolution layer, the number of input channels is 3, the number of output channels is 32, the convolution kernel size is 3*3, and the step length is 2*2; the third layer is a Relu activation layer; the fourth layer is a convolution layer, the number of input channels is 32, the number of output channels is 64, the convolution kernel size is 3*3, and the step length is 2*2; the fifth layer is a batch normalization layer; the sixth layer is a Relu activation layer; the seventh layer is a convolution layer, the number of input channels is 64, the number of output channels is 128, the convolution kernel size is 3*3, and the step length is 2*2; The slice feature map f output by the encoding submodule is expressed as follows: f=G E (x;θ e ); Among them, θ e represents the network parameters corresponding to the encoding submodule, G E (x;θ e ) represents the encoding submodule, and x represents the input training sample; The decoding submodule includes: the first layer is a deconvolution layer, the number of input channels is 128, the number of output channels is 64, the convolution kernel size is 3*3, and the step size is 2*2; the second layer is a batch normalization layer; the third layer is a Relu activation layer; the fourth layer is a deconvolution layer, the number of input channels is 64, the number of output channels is 32, the convolution kernel size is 3*3, and the step size is 2*2; the fifth layer is a batch normalization layer; the sixth layer is a Relu activation layer; the seventh layer is a deconvolution layer, the number of input channels is 32, the number of output channels is 3, the convolution kernel size is 3*3, and the step size is 2*2; Slice reconstruction of defect potential area output by the decoding submodule It is expressed as follows: Among them, θ d represents the network parameters corresponding to the decoding submodule, G D (f;θ d ) represents a decoding submodule; The re-encoding module specifically includes: Among them, θ e1 Represents the network parameters corresponding to the re-encoding module, represents the re-encoding module, Reconstructed feature map representing the output of the re-encoding module; The discrimination module specifically includes a scale one discrimination module and a scale two discrimination module, and the specific structure is as follows: The scale-one discrimination module includes: the first layer is an input layer, the number of input channels is 3; the second layer is a convolution layer, the number of input channels is 3, the number of output channels is 32, the convolution kernel size is 3*3, and the step length is 2*2; the third layer is a Relu activation layer; the fourth layer is a convolution layer, the number of input channels is 32, the number of output channels is 64, the convolution kernel size is 3*3, and the step length is 2*2; the fifth layer is a batch normalization layer; the sixth layer is a Relu activation layer; the seventh layer is a convolution layer, the number of input channels is 64, the number of output channels is 128, the convolution kernel size is 3*3, and the step length is 2*2; the eighth layer is a batch normalization layer; the ninth layer is a Relu activation layer; the tenth layer is a convolution layer, the number of input channels is 128, the number of output channels is 1, the convolution kernel size is 3*3, and the step length is 1*1; the eleventh layer is a Sigmoid activation layer; the output of the scale-one discrimination module is as follows: Among them, the result of l1 is 0 or 1. When it is 0, it means reconstructing the feature map, and when it is 1, it means slicing the image. D1 represents the network parameters corresponding to the scale-discriminant module, is the scale-discrimination module; The scale-two discrimination module includes: the first layer is an input layer with 64 channels; the second layer is a convolution layer with 64 input channels, 128 output channels, a convolution kernel size of 3*3, and a step size of 2*2; the third layer is a batch normalization layer; the fourth layer is a Relu activation layer; the fifth layer is a convolution layer with 128 input channels, 1 output channel, a convolution kernel size of 3*3, and a step size of 1*1; the sixth layer is a Sigmoid activation layer; the output of the scale-two discrimination module is as follows: l2=D2(G E-2 (x),G D-2 (f);θ D2 ); Among them, the result of l2 is 0 or 1. When it is 0, it means reconstructing the feature map, and when it is 1, it means slicing the image. D2 represents the network parameters corresponding to the scale-two discrimination module, G E-2 G represents the network module before and including the 4th layer of the encoding submodule in the generation module. D-2 Denotes the 4th layer of the decoding submodule and its previous network modules, D2(G E-2 (x),G D-2 (f);θ D2 ) is the scale-two discrimination module.
2. According to claim 1, a reverse learning two-stage intelligent recognition method for X-ray images of unequal thickness weld defects is characterized by: The method also includes marking potential defect areas of the weld of the positive sample; and performing a slicing operation on the negative sample.
3. According to claim 1, a reverse learning two-stage intelligent recognition method for X-ray images of unequal thickness weld defects is characterized in that: The preset weld defect potential area recognition model is specifically a YOLOv8 target detection network model, and parameters trained on the ImageNet dataset are used as pre-training parameters of the YOLOv8 target detection network model.
4. A reverse learning two-stage intelligent recognition system for X-ray images of unequal thickness weld defects, applied to a reverse learning two-stage intelligent recognition method for X-ray images of unequal thickness weld defects as described in any one of claims 1-3, characterized in that: include: A data set establishment module, which acquires a weld image and establishes a first data set; The image with defects in the weld image is a positive sample, and the image without defects in the weld image is a negative sample; a second data set is established based on the negative samples; the first data set includes positive samples and negative samples; A model training module, inputting the first data set into a preset weld defect potential area recognition model to train the weld defect potential area recognition model; inputting the second data set into a preset weld defect precise recognition model to train the weld defect precise recognition model; The region recognition module obtains a weld image to be recognized, inputs the weld image to be recognized into the trained weld defect potential region recognition model to recognize and obtain the defect potential region of the weld image to be recognized; slices the weld defect potential region to obtain a slice image, uses the slice image as input of the trained weld defect accurate recognition model, and recognizes and marks the slice image in which the defect potential region exists; The weld defect accurate identification model specifically includes: a generation module, a recoding module, and a discrimination module. The generation module obtains the slice feature map of the slice image and reconstructs it to obtain a slice reconstruction map; the recoding module extracts the reconstruction feature map of the slice reconstruction map; the discrimination module inputs the reconstruction feature map and the slice image, judges the reconstruction feature map and the slice image, and helps the generation module learn the characteristics of a normal weld image; The generation module specifically includes an encoding submodule and a decoding submodule, and the specific structure is as follows: The encoding submodule includes: the first layer is an input layer, the number of input channels is 3; the second layer is a convolution layer, the number of input channels is 3, the number of output channels is 32, the convolution kernel size is 3*3, and the step length is 2*2; the third layer is a Relu activation layer; the fourth layer is a convolution layer, the number of input channels is 32, the number of output channels is 64, the convolution kernel size is 3*3, and the step length is 2*2; the fifth layer is a batch normalization layer; the sixth layer is a Relu activation layer; the seventh layer is a convolution layer, the number of input channels is 64, the number of output channels is 128, the convolution kernel size is 3*3, and the step length is 2*2; The slice feature map f output by the encoding submodule is expressed as follows: f=G E (x;θ e ); Among them, θ e represents the network parameters corresponding to the encoding submodule, G E (x;θ e ) represents the encoding submodule, and x represents the input training sample; The decoding submodule includes: the first layer is a deconvolution layer, the number of input channels is 128, the number of output channels is 64, the convolution kernel size is 3*3, and the step size is 2*2; the second layer is a batch normalization layer; the third layer is a Relu activation layer; the fourth layer is a deconvolution layer, the number of input channels is 64, the number of output channels is 32, the convolution kernel size is 3*3, and the step size is 2*2; the fifth layer is a batch normalization layer; the sixth layer is a Relu activation layer; the seventh layer is a deconvolution layer, the number of input channels is 32, the number of output channels is 3, the convolution kernel size is 3*3, and the step size is 2*2; Slice reconstruction of defect potential area output by the decoding submodule It is expressed as follows: Among them, θ d represents the network parameters corresponding to the decoding submodule, G D (f;θ d ) represents a decoding submodule; The re-encoding module specifically includes: Among them, θ e1 Represents the network parameters corresponding to the re-encoding module, represents the re-encoding module, Reconstructed feature map representing the output of the re-encoding module; The discrimination module specifically includes a scale one discrimination module and a scale two discrimination module, and the specific structure is as follows: The scale-one discrimination module includes: the first layer is an input layer, the number of input channels is 3; the second layer is a convolution layer, the number of input channels is 3, the number of output channels is 32, the convolution kernel size is 3*3, and the step length is 2*2; the third layer is a Relu activation layer; the fourth layer is a convolution layer, the number of input channels is 32, the number of output channels is 64, the convolution kernel size is 3*3, and the step length is 2*2; the fifth layer is a batch normalization layer; the sixth layer is a Relu activation layer; the seventh layer is a convolution layer, the number of input channels is 64, the number of output channels is 128, the convolution kernel size is 3*3, and the step length is 2*2; the eighth layer is a batch normalization layer; the ninth layer is a Relu activation layer; the tenth layer is a convolution layer, the number of input channels is 128, the number of output channels is 1, the convolution kernel size is 3*3, and the step length is 1*1; the eleventh layer is a Sigmoid activation layer; the output of the scale-one discrimination module is as follows: Among them, the result of l1 is 0 or 1. When it is 0, it means reconstructing the feature map, and when it is 1, it means slicing the image. D1 represents the network parameters corresponding to the scale-discriminant module, is the scale-discrimination module; The scale-two discrimination module includes: the first layer is an input layer with 64 channels; the second layer is a convolution layer with 64 input channels, 128 output channels, a convolution kernel size of 3*3, and a step size of 2*2; the third layer is a batch normalization layer; the fourth layer is a Relu activation layer; the fifth layer is a convolution layer with 128 input channels, 1 output channel, a convolution kernel size of 3*3, and a step size of 1*1; the sixth layer is a Sigmoid activation layer; the output of the scale-two discrimination module is as follows: l2=D2(G E-2 (x),G D-2 (f);θ D2 ); Among them, the result of l2 is 0 or 1. When it is 0, it means reconstructing the feature map, and when it is 1, it means slicing the image. D2 represents the network parameters corresponding to the scale-two discrimination module, G E-2 G represents the network module before and including the 4th layer of the encoding submodule in the generation module. D-2 Denotes the 4th layer of the decoding submodule and its previous network modules, D2(G E-2 (x),G D-2 (f);θ D2 ) is the scale-two discrimination module.
Citation Information
Patent Citations
Two-stage mainboard image defect detecting and positioning method based on machine vision
CN114972213A