A pump body internal defect automatic identification method based on X-ray imaging
By using theoretical scattering field modeling based on Rayleigh scattering theory and an improved Swin-Unet model, combined with a Bayesian neural network, the problems of scattering artifacts and boundary blurring in X-ray imaging were solved, enabling high-precision identification and localization of internal defects in the pump body, thus improving detection efficiency and intelligence.
Patent Information
- Application Number
- CN202511251277.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-09-03
AI Technical Summary
Existing X-ray imaging methods suffer from problems such as scattering artifacts, reduced image contrast, blurred boundaries, easy obscuring or misjudgment of small defects, low identification accuracy, and insufficient positioning precision in identifying defects, especially in complex backgrounds.
We employ a theoretical scattering field modeling and inverse scattering field modeling neural network based on Rayleigh scattering theory, combined with frequency domain boundary enhancement and an improved Swin-Unet model. Through a hierarchical multi-head self-attention and symmetric hop connection encoder-decoder structure, combined with a Bayesian neural network for uncertainty modeling, we achieve high-precision defect identification and robust localization.
It achieves high-precision identification and robust positioning of internal defects in pump bodies, and has strong anti-scattering interference capability, high boundary perception capability, and quantifiable uncertainty, significantly improving detection efficiency and intelligence level.
Smart Images

Figure CN120726060B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of nondestructive testing technology, and in particular to an automatic identification method for internal defects in pump bodies based on X-ray imaging. Background Technology
[0002] With the increasing demands for internal quality of complex structural components such as pump bodies in industrial manufacturing, non-destructive testing methods based on X-ray imaging are widely used in aerospace, energy equipment, and high-end manufacturing. Existing X-ray image recognition methods mainly rely on traditional image enhancement and fixed threshold segmentation techniques to identify and locate defects, but they generally suffer from the following problems in practical applications:
[0003] Strong scattering artifacts interfere with X-ray imaging, leading to reduced image contrast, blurred boundaries, and the easy obscuring or misjudgment of minute defects. Most existing scattering correction methods are based on empirical modeling or simple image filtering, which cannot effectively recover true defect information. Traditional segmentation networks have low accuracy in defect recognition against complex backgrounds, and the limited local receptive field results in insufficient edge recognition capabilities, making it difficult to adapt to defect structures of different scales and shapes. At the same time, most methods only output a single deterministic prediction result, lacking quantification and characterization of the reliability of the recognition result, which easily leads to false detections or missed detections in areas with blurred defect boundaries or missing information. In the defect localization step, the failure to effectively combine contextual information for structured analysis results in insufficient accuracy and stability of location recognition.
[0004] Therefore, how to provide an automatic identification method for internal defects of pumps based on X-ray imaging is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] One objective of this invention is to propose an automatic identification method for internal defects of pump bodies based on X-ray imaging. This invention integrates scattering inverse modeling, frequency domain boundary enhancement, improved Swin-Unet model and Bayesian neural network inference technology to construct a fully automatic identification system from artifact correction and feature extraction to defect localization. This achieves high-precision identification and robust localization of internal defects of pump bodies, and has the advantages of strong anti-scattering interference capability, high boundary perception capability, quantifiable uncertainty and high localization accuracy, thereby improving detection efficiency and intelligence level.
[0006] An automatic identification method for internal defects of a pump body based on X-ray imaging according to an embodiment of the present invention includes the following steps:
[0007] Step 1: Acquire raw X-ray images of the inside of the pump body under test and obtain pump body parameters;
[0008] Step 2: Based on the pump parameters and Rayleigh scattering theory, generate a theoretical scattering field image;
[0009] Step 3: Input the original X-ray image and the theoretical scattering field image into the scattering field inverse modeling neural network, and output the scattering artifact correction image;
[0010] Step 4: Perform frequency domain decomposition on the scattering artifact corrected image, perform boundary enhancement processing, and generate a boundary enhanced image;
[0011] Step 5: Input the boundary enhancement image into the improved Swin-Unet model. The improvement of the improved Swin-Unet model is that it adopts a hierarchical multi-head self-attention module and a symmetric skip connection encoder-decoder structure to output a defect feature map.
[0012] Step 6: Perform spatial dimension decoding on the defect feature map to generate a pixel-level defect prediction image;
[0013] Step 7: Input the pixel-level defect prediction image into a Bayesian neural network and perform multiple forward samplings to generate an uncertainty image;
[0014] Step 8: Based on the pixel-level defect prediction image and the uncertainty image, calculate the comprehensive confidence score of each pixel, and perform connected component analysis to identify the location coordinates of the defect.
[0015] Optionally, the pump body parameters specifically include the material type, density, and geometric parameters of the pump body.
[0016] Optionally, step two specifically includes:
[0017] A scattering voxel model based on the pump body geometry parameters is constructed in three-dimensional space. Within each voxel unit, the density and effective atomic number of the material corresponding to the voxel unit are obtained. Combined with the spatial direction from the center of the voxel unit to the X-ray source, the incident angle of the spatial direction is determined.
[0018] Based on the incident angle and material properties, Rayleigh scattering theory is used to discretely sample multiple scattering angles within a preset range for each voxel unit to obtain the scattering intensity value at each scattering angle.
[0019] The scattering intensity values at each scattering angle are weighted according to the angle between the scattering direction and the normal of the detection plane to obtain a weighted scattering intensity value. The weighted scattering intensity values are then summed to obtain the effective scattering contribution value of the voxel unit in the imaging direction.
[0020] In the scattering voxel model, the effective scattering contribution of all voxel units along the X-ray incident direction is integrated and accumulated to form a three-dimensional intensity distribution map;
[0021] The three-dimensional intensity distribution map is transformed by two-dimensional projection, and the cumulative scattering value of each point is mapped to the corresponding pixel point by combining the positional relationship of the imaging plane to form a theoretical scattering field image.
[0022] Optionally, step three specifically involves:
[0023] The original X-ray image and the theoretical scattering field image are stitched together along the channel dimension to form the input image;
[0024] The input image is fed into a scattering field inverse modeling neural network. The scattering field inverse modeling neural network adopts an encoder-decoder structure. The encoder includes a four-level feature extraction module, and each level of the feature extraction module includes:
[0025] A two-dimensional convolutional layer with a kernel size of 3×3, a stride of 1, and padding of 1;
[0026] A batch normalization layer;
[0027] A ReLU activation function layer;
[0028] A max pooling layer with a kernel size of 2×2 and a stride of 2;
[0029] The decoder includes four upsampling modules, each of which includes:
[0030] A deconvolutional layer with a kernel size of 4×4 and a stride of 2 is used to obtain an upsampled image.
[0031] Channel stitching operation is used to generate a feature stitched image;
[0032] Two 3×3 convolutional layers, each followed by a batch normalization layer and a ReLU activation function layer, are used to receive the feature stitched map and obtain the fused feature map of the current level;
[0033] The channel stitching operation specifically includes: the first-level upsampling module stitches the upsampled image output from the deconvolution layer with the output feature image from the third-level feature extraction module of the encoder; the second-level upsampling module stitches the upsampled image output from the deconvolution layer with the output feature image from the second-level feature extraction module of the encoder; the third-level upsampling module stitches the upsampled image output from the deconvolution layer with the output feature image from the first-level feature extraction module of the encoder; the fourth-level upsampling module does not stitch the encoder output, but only performs upsampling processing.
[0034] The fused feature map output by each upsampling module in the decoder is used as the input of the next upsampling module. The input of the first upsampling module is the output feature map of the fourth feature extraction module of the encoder. They are connected layer by layer in sequence to form a decoding path.
[0035] After the fourth-level upsampling module of the decoder, a two-dimensional convolutional layer with a kernel size of 1×1 is connected to compress the number of channels of the fused feature map output by the fourth-level upsampling module to a single channel and output the residual image.
[0036] The residual image is subtracted from the original X-ray image pixel by pixel to generate a scattering artifact corrected image.
[0037] Optionally, step four specifically involves:
[0038] Perform a two-dimensional fast Fourier transform on the scattering artifact-corrected image to obtain the frequency domain amplitude spectrum;
[0039] A rectangular bandpass mask is set in the frequency domain amplitude spectrum. The rectangular bandpass mask only retains the frequency components whose center is the origin and whose radius is within the range of 20% to 60% of the normalized frequency. The amplitude values corresponding to the other frequency positions are set to zero.
[0040] The rectangular bandpass mask is applied to the frequency domain amplitude spectrum to obtain high-frequency edge information, and a two-dimensional inverse fast Fourier transform is performed to obtain an edge response map containing edge details.
[0041] The edge response map is subjected to pixel intensity normalization processing, and the pixel values are linearly scaled to between 0 and 1;
[0042] The normalized edge response map and the scattering artifact correction image are weighted and superimposed according to a preset ratio to generate a boundary enhancement image.
[0043] Optionally, step five specifically includes:
[0044] The boundary enhancement image is input into the improved Swin-Unet model and divided into 4×4 image blocks. Patch embedding is completed through a convolution operation with a stride of 4, which is mapped to a token feature map of a set dimension.
[0045] The encoder includes a four-level Swing Transformer encoding module. Each Swing Transformer encoding module adopts a hierarchical multi-head self-attention module, specifically including: a non-overlapping window multi-head self-attention module and a sliding window multi-head self-attention module. The non-overlapping window multi-head self-attention module and the sliding window multi-head self-attention module are stacked alternately.
[0046] The non-overlapping window module divides the input Token feature map into fixed-size windows and performs a multi-head attention mechanism independently within each window. By performing linear mapping and weighted summation on the query, key, and value vectors of the Token feature map within the window, it captures the feature correlation within the local region.
[0047] The sliding window multi-head self-attention module re-divides the input Token feature map into windows by offsetting the non-overlapping window of the previous layer by half the window size, thereby achieving cross-coverage between windows and performing multi-head self-attention calculation within the new window;
[0048] Each Swin Transformer encoding module is connected through a Patch Merging module, which performs spatial downsampling and feature compression by concatenating adjacent token feature maps and performing linear transformation.
[0049] The decoder includes a four-level Patch Expanding module for progressively upsampling to restore spatial resolution. The outputs of each level of the decoder are spliced and fused with the token feature maps of the corresponding level outputs in the encoder through symmetrical skip connections.
[0050] The decoder is connected to a 1×1 two-dimensional convolutional layer at the end, which is used to perform single-channel compression on the final fused feature map and output a defect feature map.
[0051] Optionally, step six specifically includes:
[0052] The defect feature map is upsampled three times by bilinear interpolation, and the image size is magnified by two times each time so that the output image is restored to the same spatial resolution as the boundary enhancement image.
[0053] The upsampled defect feature map is input into a 2D convolutional layer with a kernel size of 3×3 to refine the edge features. The pixel values are then mapped to between 0 and 1 using the Sigmoid activation function to generate a pixel-level defect prediction image.
[0054] Optionally, step seven specifically includes:
[0055] The Bayesian neural network inserts a Dropout structure into each hidden layer, sets a Dropout rate to randomly discard feature channels, and keeps Dropout active during the inference phase. It repeats forward propagation N times for the same input pixel-level defect prediction image to generate N independent sub-prediction images.
[0056] Obtain the predicted values at each pixel of all sub-prediction images, and summarize them into multiple sets of predicted values for each pixel.
[0057] For each pixel, the mean of all predicted values in the predicted value set is obtained, the squared difference between all predicted values and the mean is calculated, and the average of all squared differences is obtained to get the prediction variance value of the pixel.
[0058] The prediction variance of each pixel is used as the corresponding uncertainty metric to ultimately form an uncertainty image.
[0059] Optionally, step eight specifically includes:
[0060] Based on the pixel-level defect prediction image and the uncertainty image, a comprehensive confidence score is calculated for each pixel. The comprehensive confidence score is obtained by jointly calculating the pixel value of the pixel-level defect prediction image and the corresponding uncertainty metric.
[0061] For each pixel, if the overall confidence score is higher than the preset score threshold, the pixel is assigned the value 1; otherwise, the pixel is assigned the value 0, resulting in a binary image.
[0062] Connectivity analysis is performed on the binary image. Each connected component is determined using a four-adjacency method. The bounding rectangle of each connected component is extracted based on its spatial coverage. The center coordinates of the bounding rectangle of each connected component are then used as the location coordinates of the defect.
[0063] The beneficial effects of this invention are:
[0064] This invention addresses the problem of high-intensity scattering interference caused by complex structural materials in X-ray images by constructing a theoretical scattering field modeling and inverse scattering field modeling neural network based on Rayleigh scattering theory. In the frequency domain processing stage, a rectangular bandpass mask is introduced to obtain multi-scale edge response maps, which are then weighted and superimposed with the original image to enhance image details and suppress background. To address the limitations of existing network structures in extracting minute defects, the Swin-Unet model structure is improved by fusing a hierarchical multi-head self-attention mechanism with a sliding window strategy to enhance feature representation capabilities, and symmetrical skip connections are combined to improve the efficiency of contextual information transmission. A Bayesian neural network and Monte Carlo Dropout strategy are introduced to model the uncertainty of pixel-level defect prediction results, improving the reliability of the model's identification results in complex and blurred regions. Finally, by combining comprehensive confidence score filtering with four-adjacent connected component analysis, the spatial location of the defect region is accurately located, outputting high-resolution positioning coordinates. This invention achieves high-accuracy, high-confidence automatic identification and intelligent positioning of fine-grained defects in pump X-ray images, significantly improving the stability and interpretability of defect detection. Attached Figure Description
[0065] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0066] Figure 1This is an overall flowchart of an automatic identification method for internal defects of a pump body based on X-ray imaging, as proposed in this invention.
[0067] Figure 2 This is a flowchart of the boundary enhancement process for an automatic identification method for internal defects of a pump body based on X-ray imaging proposed in this invention.
[0068] Figure 3 This is a flowchart illustrating the process of extracting defect location coordinates using connected component analysis in an automatic identification method for internal defects of a pump body based on X-ray imaging, as proposed in this invention. Detailed Implementation
[0069] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0070] refer to Figures 1-3 An automatic identification method for internal defects of pump bodies based on X-ray imaging includes the following steps:
[0071] Step 1: Acquire raw X-ray images of the inside of the pump body under test and obtain pump body parameters;
[0072] Step 2: Based on the pump parameters and Rayleigh scattering theory, generate a theoretical scattering field image;
[0073] Step 3: Input the original X-ray image and the theoretical scattering field image into the scattering field inverse modeling neural network, and output the scattering artifact correction image;
[0074] Step 4: Perform frequency domain decomposition on the scattering artifact corrected image, perform boundary enhancement processing, and generate a boundary enhanced image;
[0075] Step 5: Input the boundary enhancement image into the improved Swin-Unet model. The improvement of the improved Swin-Unet model is that it adopts a hierarchical multi-head self-attention module and a symmetric skip connection encoder-decoder structure to output a defect feature map.
[0076] Step 6: Perform spatial dimension decoding on the defect feature map to generate a pixel-level defect prediction image;
[0077] Step 7: Input the pixel-level defect prediction image into a Bayesian neural network and perform multiple forward samplings to generate an uncertainty image;
[0078] Step 8: Based on the pixel-level defect prediction image and the uncertainty image, calculate the comprehensive confidence score of each pixel, and perform connected component analysis to identify the location coordinates of the defect.
[0079] In this embodiment, the pump body parameters specifically include the material type, density, and geometric parameters of the pump body.
[0080] In this embodiment, step two specifically includes:
[0081] A scattering voxel model based on the pump body geometry parameters is constructed in three-dimensional space. Within each voxel unit, the density and effective atomic number of the material corresponding to the voxel unit are obtained. Combined with the spatial direction from the center of the voxel unit to the X-ray source, the incident angle of the spatial direction is determined.
[0082] Based on the incident angle and material properties, Rayleigh scattering theory is used to discretely sample multiple scattering angles within a preset range for each voxel unit to obtain the scattering intensity value at each scattering angle.
[0083] ;
[0084] in, Indicates the first The scattering intensity value at each scattering angle. Indicates the intensity of incident X-rays. Indicates the effective atomic number of the material. Indicates wavelength. Indicates the material density adjustment factor;
[0085] The scattering intensity values at each scattering angle are weighted according to the angle between the scattering direction and the normal of the detection plane to obtain a weighted scattering intensity value. The weighted scattering intensity values are then summed to obtain the effective scattering contribution value of the voxel unit in the imaging direction.
[0086] ;
[0087] in, Indicates the first The weighted scattering intensity value at each scattering angle. This represents the angle between the scattering direction and the normal to the detection plane;
[0088] In the scattering voxel model, the effective scattering contribution of all voxel units along the X-ray incident direction is integrated and accumulated to form a three-dimensional intensity distribution map;
[0089] The three-dimensional intensity distribution map is transformed by two-dimensional projection, and the cumulative scattering value of each point is mapped to the corresponding pixel point by combining the positional relationship of the imaging plane to form a theoretical scattering field image.
[0090] In this embodiment, step three specifically involves:
[0091] The original X-ray image and the theoretical scattering field image are stitched together along the channel dimension to form the input image;
[0092] The input image is fed into a scattering field inverse modeling neural network. The scattering field inverse modeling neural network adopts an encoder-decoder structure. The encoder includes a four-level feature extraction module, and each level of the feature extraction module includes:
[0093] A two-dimensional convolutional layer with a kernel size of 3×3, a stride of 1, and padding of 1;
[0094] A batch normalization layer;
[0095] A ReLU activation function layer;
[0096] A max pooling layer with a kernel size of 2×2 and a stride of 2;
[0097] The decoder includes four upsampling modules, each of which includes:
[0098] A deconvolutional layer with a kernel size of 4×4 and a stride of 2 is used to obtain an upsampled image.
[0099] Channel stitching operation is used to generate a feature stitched image;
[0100] Two 3×3 convolutional layers, each followed by a batch normalization layer and a ReLU activation function layer, are used to receive the feature stitched map and obtain the fused feature map of the current level;
[0101] The channel stitching operation specifically includes: the first-level upsampling module stitches the upsampled image output from the deconvolution layer with the output feature image from the third-level feature extraction module of the encoder; the second-level upsampling module stitches the upsampled image output from the deconvolution layer with the output feature image from the second-level feature extraction module of the encoder; the third-level upsampling module stitches the upsampled image output from the deconvolution layer with the output feature image from the first-level feature extraction module of the encoder; the fourth-level upsampling module does not stitch the encoder output, but only performs upsampling processing.
[0102] The fused feature map output by each upsampling module in the decoder is used as the input of the next upsampling module. The input of the first upsampling module is the output feature map of the fourth feature extraction module of the encoder. They are connected layer by layer in sequence to form a decoding path.
[0103] After the fourth-level upsampling module of the decoder, a two-dimensional convolutional layer with a kernel size of 1×1 is connected to compress the number of channels of the fused feature map output by the fourth-level upsampling module to a single channel and output the residual image.
[0104] The residual image is subtracted from the original X-ray image pixel by pixel to generate a scattering artifact corrected image.
[0105] In this invention, the scattering field inverse modeling neural network adopts a symmetrical encoder-decoder structure comprising a four-level encoder and a four-level decoder. The encoder section employs a four-level progressive convolution and downsampling structure, enabling the extraction of spatial semantic features at different scales of the image layer by layer. The decoder adopts a symmetrical structure design and introduces multi-level skip connections to effectively alleviate the information loss problem during feature decoding. Through channel concatenation operations, features extracted from different levels of the encoder are directly introduced into the corresponding decoder modules, allowing the corrected image to accurately restore fine-grained edge and texture information while preserving global structural consistency.
[0106] In this embodiment, step four specifically includes:
[0107] Perform a two-dimensional fast Fourier transform on the scattering artifact-corrected image to obtain the frequency domain amplitude spectrum;
[0108] A rectangular bandpass mask is set in the frequency domain amplitude spectrum. The rectangular bandpass mask only retains the frequency components whose center is the origin and whose radius is within the range of 20% to 60% of the normalized frequency. The amplitude values corresponding to the other frequency positions are set to zero.
[0109] The rectangular bandpass mask is applied to the frequency domain amplitude spectrum to obtain high-frequency edge information, and a two-dimensional inverse fast Fourier transform is performed to obtain an edge response map containing edge details.
[0110] The edge response map is subjected to pixel intensity normalization processing, and the pixel values are linearly scaled to between 0 and 1;
[0111] The normalized edge response map and the scattering artifact correction image are weighted and superimposed according to a preset ratio to generate a boundary enhancement image.
[0112] In this embodiment, step five specifically includes:
[0113] The boundary enhancement image is input into the improved Swin-Unet model and divided into 4×4 image blocks. Patch embedding is completed through a convolution operation with a stride of 4, which is mapped to a token feature map of a set dimension.
[0114] The encoder includes a four-level Swing Transformer encoding module. Each Swing Transformer encoding module adopts a hierarchical multi-head self-attention module, specifically including: a non-overlapping window multi-head self-attention module and a sliding window multi-head self-attention module. The non-overlapping window multi-head self-attention module and the sliding window multi-head self-attention module are superimposed in an alternating manner to enhance the local and global dependency modeling capabilities.
[0115] The non-overlapping window module divides the input Token feature map into fixed-size windows and performs a multi-head attention mechanism independently within each window. By performing linear mapping and weighted summation on the query, key, and value vectors of the Token feature map within the window, it captures the feature correlation within the local region.
[0116] The sliding window multi-head self-attention module re-divides the input Token feature map into windows by offsetting it by half the window size from the non-overlapping window of the previous layer, thereby achieving cross-coverage between windows. Multi-head self-attention calculation is performed within the new window, thereby introducing cross-window feature interaction and effectively improving the spatial modeling range and context information perception capability.
[0117] Each Swin Transformer encoding module is connected through a Patch Merging module, which performs spatial downsampling and feature compression by concatenating adjacent token feature maps and performing linear transformation.
[0118] The decoder includes a four-level Patch Expanding module for progressively upsampling to restore spatial resolution. The outputs of each level of the decoder are spliced and fused with the token feature maps of the corresponding level outputs in the encoder through symmetrical skip connections.
[0119] The decoder is connected to a 1×1 two-dimensional convolutional layer at the end, which is used to perform single-channel compression on the final fused feature map and output a defect feature map.
[0120] In this invention, to address the challenges of identifying defects in X-ray images, such as blurred boundaries, subtle local changes, and strong contextual dependencies, an improved Swin-Unet network model is designed. This model introduces multi-head self-attention modules with non-overlapping windows and sliding windows in the encoder, stacked alternately. The non-overlapping window module divides the image into fixed windows, performing attention calculations within each window to model the fine features of the local region. The sliding window module re-divides the previous layer's window by offsetting it by half a window size, enabling cross-window feature interaction, expanding the attention receptive field, and effectively enhancing spatial modeling capabilities and contextual information perception.
[0121] The patch embedding stage replaces traditional linear projection with a convolution operation with a stride of 4, mapping image patches to token feature maps of a set dimension. This enhances the ability to express local edges and structural details, improving the limitations of the original embedding in fine-grained feature extraction. Between encoder layers, the patch merging module achieves spatial downsampling and feature compression, maintaining the semantic integrity of the multi-scale token structure. The decoder consists of multi-level patch expanding modules, progressively upsampling to restore spatial resolution, and using symmetrical skip connections to concatenate and fuse the encoded and decoded feature maps at each level, achieving the integration of information at different levels and the restoration of details.
[0122] The improved Swin-Unet model offers more targeted regional modeling for minute defects in images, significantly enhancing both feature representation and spatial resolution. The model can accurately extract defect edges and morphological structures even with boundary-enhanced image input, demonstrating good detection accuracy and adaptability, reflecting substantial technological advancements.
[0123] In this embodiment, step six specifically includes:
[0124] The defect feature map is upsampled three times by bilinear interpolation, and the image size is magnified by two times each time so that the output image is restored to the same spatial resolution as the boundary enhancement image.
[0125] The upsampled defect feature map is input into a two-dimensional convolutional layer with a kernel size of 3×3 to refine the edge features. The pixel values are then mapped to between 0 and 1 using the Sigmoid activation function to generate a pixel-level defect prediction image.
[0126] Because the improved Swin-Unet model employs a multi-level Patch Merging structure to achieve layer-by-layer spatial downsampling, it extracts multi-scale deep features through continuous hierarchical nested modules in the encoder stage, resulting in a significant reduction in the spatial resolution of the output defect feature map compared to the input image. Although the decoder uses the Patch Expanding module to upsample layer by layer to recover spatial information, in order to balance model computational efficiency and feature representation capability, the final generated defect feature map may still not be completely restored to the same resolution as the original image. Therefore, a spatial dimension decoding operation is introduced in step six above to further upsample the defect feature map, so that its output meets the pixel-by-pixel localization requirements, generating a pixel-level defect prediction image with the same size as the boundary enhancement image.
[0127] In this embodiment, step seven specifically includes:
[0128] The Bayesian neural network inserts a Dropout structure into each hidden layer, sets a Dropout rate to randomly discard feature channels, and keeps Dropout active during the inference phase. It repeats forward propagation N times for the same input pixel-level defect prediction image to generate N independent sub-prediction images.
[0129] In this embodiment, by inserting Dropout structures into each hidden layer of the Bayesian neural network and setting a fixed Dropout rate, the feature channels of each layer are randomly dropped, allowing the Bayesian neural network to form different neuron activation paths in each forward propagation. This strategy essentially constructs a set of subnetworks with shared weights but randomly varying substructures, thereby simulating the probability distribution characteristics of model parameters without introducing additional training overhead, thus improving the network's robustness and generalization ability to input perturbations. By keeping the Dropout structures active during the inference phase, different predicted outputs can be obtained in multiple forward propagations, further providing a sufficient statistical sample basis for uncertainty estimation.
[0130] Obtain the predicted values at each pixel of all sub-prediction images, and summarize them into multiple sets of predicted values for each pixel.
[0131] For each pixel, the mean of all predicted values in the predicted value set is obtained, the squared difference between all predicted values and the mean is calculated, and the average of all squared differences is obtained to get the prediction variance value of the pixel.
[0132] The prediction variance of each pixel is used as the corresponding uncertainty metric to ultimately form an uncertainty image.
[0133] In this embodiment, step eight specifically includes:
[0134] Based on the pixel-level defect prediction image and the uncertainty image, a comprehensive confidence score is calculated for each pixel. The comprehensive confidence score is obtained by jointly calculating the pixel value of the pixel-level defect prediction image and the corresponding uncertainty metric.
[0135] ;
[0136] in, Represents pixels The overall confidence score, Represents pixels in a pixel-level defect prediction image pixel values, Representing pixels in an image with uncertainty Uncertainty metrics The maximum uncertainty metric in the uncertainty image;
[0137] For each pixel, if the overall confidence score is higher than the preset score threshold, the pixel is assigned the value 1; otherwise, the pixel is assigned the value 0, resulting in a binary image.
[0138] Connectivity analysis is performed on the binary image. Each connected component is determined using the four-adjacency method. The bounding rectangle of each connected component is extracted based on its spatial coverage. The center coordinates of the bounding rectangle of each connected component are then used as the location coordinates of the defect.
[0139] When performing connected component analysis, the four-adjacency approach means that each pixel is only checked for connectivity with its four adjacent pixels in the image (up, down, left, and right). The entire binary image is traversed pixel by pixel. For each pixel scanned, if its value is 1, its four adjacent pixels (up, down, left, and right) are checked to see if they are also 1 and not yet included in a connected component. If these conditions are met, the adjacent pixel is marked as part of the current connected component, and it becomes the new starting point for adjacency expansion search. This process is recursively repeated or advanced using a queue until all reachable four-adjacent pixels in the current connected component have been traversed and marked, forming a complete connected component.
[0140] After marking a connected component, the process continues to search for the next unvisited pixel with a value of 1 in the entire image, using it as the starting point for a new connected component. This process is repeated until all foreground pixels (pixels with a value of 1) in the entire binary image are assigned to a connected component, thus achieving the segmentation and recognition of all connected regions in the binary image. The four-adjacency method is suitable for scenarios where the defect region has a clear structural boundary, avoiding misjudging diagonally adjacent pixels as the same region.
[0141] Example 1:
[0142] To verify the feasibility of this invention in practice, it was applied to an industrial non-destructive testing production line to automatically identify internal defects in a batch of centrifugal pump bodies with complex structures and dense materials. These pump bodies are characterized by thick walls and multiple cavities, making them highly susceptible to multiple scattering artifacts and uneven imaging contrast during X-ray imaging. Traditional manual identification methods or threshold-based image segmentation techniques perform poorly in terms of accuracy and consistency, exhibiting significant problems of high false negative rates and low operational efficiency.
[0143] In this embodiment, the original X-ray image of the pump body under test is first obtained using a standard X-ray detection device. The theoretical scattering field image is generated using Rayleigh scattering theory. Then, the original X-ray image and the theoretical scattering field image are input into the scattering field inverse modeling neural network to obtain a scattering artifact corrected image. Compared with the original image, which has blurred boundaries and uneven brightness, the corrected image significantly enhances the visibility of the defect area and the clarity of the structural edges.
[0144] After performing a frequency domain transformation on the scattering artifact-corrected image, a rectangular bandpass mask is used to extract 20% to 60% of the mid-frequency components to highlight edge details, resulting in an edge response map. This map is then weighted and fused with the original image to form a boundary enhancement image. This boundary enhancement image is then input into an improved Swin-Unet model. This model, by introducing a hierarchical multi-head self-attention mechanism and a symmetric skip connection structure, significantly improves the model's ability to perceive minute defects and its spatial feature recovery performance. The output defect feature map is upsampled and activated with a sigmoid function to generate a pixel-level defect prediction image.
[0145] To further improve the reliability of the detection results, this embodiment utilizes a Bayesian neural network and Dropout strategy to perform 20 forward samplings on the same input image, generating a set of sub-predicted images, and calculating the variance of the predicted value for each pixel as an uncertainty index. Based on this, a pixel-level comprehensive confidence score is calculated by combining the pixel predicted value and the uncertainty index. The image is binarized by setting a score threshold to obtain candidate defect region images. Subsequently, connected component analysis is performed using a four-neighbor approach. By traversing and classifying each connected pixel and extracting its bounding rectangle, the center coordinates of each defect are finally output as the precise location.
[0146] Table 1. Performance Comparison of the Invention Method and Typical Detection Algorithms (500 Sample Images)
[0147]
[0148] Table 1 shows the performance comparison of three different detection models on 500 X-ray sample images, including the traditional U-Net model, the DeepLabV3+ model, and the improved Swin-Unet model proposed in this invention. The evaluation metrics cover average precision, average recall, F1 score, mean IoU, false positives, false negatives, and average detection time, comprehensively reflecting the actual effect of each algorithm in the automatic identification of defects inside the pump body.
[0149] In terms of accuracy, the traditional U-Net model achieved an average precision of 81.2%, a recall of 75.6%, an F1 score of 0.78, and a mean IoU of 0.65, placing it at a level that is basically usable but with limited accuracy. Its false positives and false negatives were 47 and 63 respectively, indicating a weak ability to detect subtle defects and a certain degree of false alarms. The detection time was 3.75 seconds, meeting medium-speed requirements.
[0150] In comparison, the DeepLabV3+ model achieved performance improvements across multiple metrics. Average precision increased to 84.9%, recall to 79.3%, F1 score to 0.82, and mean IoU to 0.70, indicating that this model outperforms traditional methods in feature extraction and context modeling. Its false positives and false negatives decreased to 38 and 51 respectively, while the average detection time slightly increased to 4.10 seconds, showing that some detection speed was sacrificed while improving accuracy.
[0151] The improved Swin-Unet model proposed in this invention outperforms the other two models across all evaluation dimensions, achieving an average precision of 91.4%, an average recall of 88.2%, an F1 score of 0.89, and a mean IoU of 0.81, comprehensively surpassing the previous two models. In particular, it reduces false positives and false negatives to 11 and 14 respectively, demonstrating exceptional recognition stability and small target detection capabilities. Furthermore, the improved Swin-Unet model has an average detection time of 2.90 seconds, the fastest among the three, showcasing its superior real-time performance and engineering application value.
[0152] This embodiment combines an improved Swin-Unet model with a Bayesian uncertainty estimation method to construct a high-precision and robust automatic defect identification process for pump body X-ray images, significantly improving the ability to identify minute defect regions in complex backgrounds. Simultaneously, the introduction of an uncertainty quantification strategy effectively enhances the reliability assessment capability of the results, and in the post-processing stage, precise defect location is achieved through comprehensive confidence analysis and connected component extraction. This invention not only demonstrates excellent recognition accuracy and stability but also possesses good processing efficiency and practical value, making it suitable for the intelligent defect identification needs in non-destructive testing of industrial products.
[0153] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for automatic identification of internal defects in a pump body based on X-ray imaging, characterized in that, Includes the following steps: Step 1: Acquire raw X-ray images of the inside of the pump body under test and obtain pump body parameters; Step 2: Based on the pump parameters and Rayleigh scattering theory, generate a theoretical scattering field image; Step 3: Input the original X-ray image and the theoretical scattering field image into the scattering field inverse modeling neural network, and output the scattering artifact correction image; Step 4: Perform frequency domain decomposition on the scattering artifact corrected image, perform boundary enhancement processing, and generate a boundary enhanced image; Step 5: Input the boundary enhancement image into the improved Swin-Unet model. The improvement of the improved Swin-Unet model is that it adopts a hierarchical multi-head self-attention module and a symmetric skip connection encoder-decoder structure to output a defect feature map. Step 6: Perform spatial dimension decoding on the defect feature map to generate a pixel-level defect prediction image; Step 7: Input the pixel-level defect prediction image into a Bayesian neural network and perform multiple forward samplings to generate an uncertainty image; Step 8: Based on the pixel-level defect prediction image and the uncertainty image, calculate the comprehensive confidence score of each pixel, and perform connected component analysis to identify the location coordinates of the defect.
2. The method for automatic identification of internal defects in a pump body based on X-ray imaging according to claim 1, characterized in that, The pump body parameters specifically include the material type, density, and geometric parameters of the pump body.
3. The method for automatic identification of internal defects in a pump body based on X-ray imaging according to claim 1, characterized in that, Step two specifically involves: A scattering voxel model based on the pump body geometry parameters is constructed in three-dimensional space. Within each voxel unit, the density and effective atomic number of the material corresponding to the voxel unit are obtained. Combined with the spatial direction from the center of the voxel unit to the X-ray source, the incident angle of the spatial direction is determined. Based on the incident angle and material properties, Rayleigh scattering theory is used to discretely sample multiple scattering angles within a preset range for each voxel unit to obtain the scattering intensity value at each scattering angle. The scattering intensity values at each scattering angle are weighted according to the angle between the scattering direction and the normal of the detection plane to obtain a weighted scattering intensity value. The weighted scattering intensity values are then summed to obtain the effective scattering contribution value of the voxel unit in the imaging direction. In the scattering voxel model, the effective scattering contribution of all voxel units along the X-ray incident direction is integrated and accumulated to form a three-dimensional intensity distribution map; The three-dimensional intensity distribution map is transformed by two-dimensional projection, and the cumulative scattering value of each point is mapped to the corresponding pixel point by combining the positional relationship of the imaging plane to form a theoretical scattering field image.
4. The method for automatic identification of internal defects in a pump body based on X-ray imaging according to claim 1, characterized in that, Step three specifically involves: The original X-ray image and the theoretical scattering field image are stitched together along the channel dimension to form the input image; The input image is fed into a scattering field inverse modeling neural network. The scattering field inverse modeling neural network adopts an encoder-decoder structure. The encoder includes a four-level feature extraction module, and each level of the feature extraction module includes: A two-dimensional convolutional layer with a kernel size of 3×3, a stride of 1, and padding of 1; A batch normalization layer; A ReLU activation function layer; A max pooling layer with a kernel size of 2×2 and a stride of 2; The decoder includes four upsampling modules, each of which includes: A deconvolutional layer with a kernel size of 4×4 and a stride of 2 is used to obtain an upsampled image. Channel stitching operation is used to generate a feature stitched image; Two 3×3 convolutional layers, each followed by a batch normalization layer and a ReLU activation function layer, are used to receive the feature stitched map and obtain the fused feature map of the current level; The channel stitching operation specifically includes: the first-level upsampling module stitches the upsampled image output from the deconvolution layer with the output feature image from the third-level feature extraction module of the encoder; the second-level upsampling module stitches the upsampled image output from the deconvolution layer with the output feature image from the second-level feature extraction module of the encoder; the third-level upsampling module stitches the upsampled image output from the deconvolution layer with the output feature image from the first-level feature extraction module of the encoder; the fourth-level upsampling module does not stitch the encoder output, but only performs upsampling processing. The fused feature map output by each upsampling module in the decoder is used as the input of the next upsampling module. The input of the first upsampling module is the output feature map of the fourth feature extraction module of the encoder. They are connected layer by layer in sequence to form a decoding path. After the fourth-level upsampling module of the decoder, a two-dimensional convolutional layer with a kernel size of 1×1 is connected to compress the number of channels of the fused feature map output by the fourth-level upsampling module to a single channel and output the residual image. The residual image is subtracted from the original X-ray image pixel by pixel to generate a scattering artifact corrected image.
5. The method for automatic identification of internal defects in a pump body based on X-ray imaging according to claim 1, characterized in that, Step four specifically involves: Perform a two-dimensional fast Fourier transform on the scattering artifact-corrected image to obtain the frequency domain amplitude spectrum; A rectangular bandpass mask is set in the frequency domain amplitude spectrum. The rectangular bandpass mask only retains the frequency components whose center is the origin and whose radius is within the range of 20% to 60% of the normalized frequency. The amplitude values corresponding to the other frequency positions are set to zero. The rectangular bandpass mask is applied to the frequency domain amplitude spectrum to obtain high-frequency edge information, and a two-dimensional inverse fast Fourier transform is performed to obtain an edge response map containing edge details. The edge response map is subjected to pixel intensity normalization processing, and the pixel values are linearly scaled to between 0 and 1; The normalized edge response map and the scattering artifact correction image are weighted and superimposed according to a preset ratio to generate a boundary enhancement image.
6. The method for automatic identification of internal defects in a pump body based on X-ray imaging according to claim 1, characterized in that, Step five specifically involves: The boundary enhancement image is input into the improved Swin-Unet model and divided into 4×4 image blocks. Patch embedding is completed through a convolution operation with a stride of 4, which is mapped to a token feature map of a set dimension. The encoder includes a four-level Swing Transformer encoding module. Each Swing Transformer encoding module adopts a hierarchical multi-head self-attention module, specifically including: a non-overlapping window multi-head self-attention module and a sliding window multi-head self-attention module. The non-overlapping window multi-head self-attention module and the sliding window multi-head self-attention module are stacked alternately. The non-overlapping window module divides the input Token feature map into fixed-size windows and performs a multi-head attention mechanism independently within each window. By performing linear mapping and weighted summation on the query, key, and value vectors of the Token feature map within the window, it captures the feature correlation within the local region. The sliding window multi-head self-attention module re-divides the input Token feature map into windows by offsetting the non-overlapping window of the previous layer by half the window size, thereby achieving cross-coverage between windows and performing multi-head self-attention calculation within the new window; Each Swin Transformer encoding module is connected through a Patch Merging module, which performs spatial downsampling and feature compression by concatenating adjacent token feature maps and performing linear transformation. The decoder includes a four-level Patch Expanding module for progressively upsampling to restore spatial resolution. The outputs of each level of the decoder are spliced and fused with the token feature maps of the corresponding level outputs in the encoder through symmetrical skip connections. The decoder is connected to a 1×1 two-dimensional convolutional layer at the end, which is used to perform single-channel compression on the final fused feature map and output a defect feature map.
7. The method for automatic identification of internal defects in a pump body based on X-ray imaging according to claim 1, characterized in that, Step six specifically involves: The defect feature map is upsampled three times by bilinear interpolation, and the image size is magnified by two times each time so that the output image is restored to the same spatial resolution as the boundary enhancement image. The upsampled defect feature map is input into a 2D convolutional layer with a kernel size of 3×3 to refine the edge features. The pixel values are then mapped to between 0 and 1 using the Sigmoid activation function to generate a pixel-level defect prediction image.
8. The method for automatic identification of internal defects in a pump body based on X-ray imaging according to claim 1, characterized in that, Step seven specifically involves: The Bayesian neural network inserts a Dropout structure into each hidden layer, sets a Dropout rate to randomly discard feature channels, and keeps Dropout active during the inference phase. It repeats forward propagation N times for the same input pixel-level defect prediction image to generate N independent sub-prediction images. Obtain the predicted values at each pixel of all sub-prediction images, and summarize them into multiple sets of predicted values for each pixel. For each pixel, the mean of all predicted values in the predicted value set is obtained, the squared difference between all predicted values and the mean is calculated, and the average of all squared differences is obtained to get the prediction variance value of the pixel. The prediction variance of each pixel is used as the corresponding uncertainty metric to ultimately form an uncertainty image.
9. The method for automatic identification of internal defects in a pump body based on X-ray imaging according to claim 1, characterized in that, Step eight specifically involves: Based on the pixel-level defect prediction image and the uncertainty image, a comprehensive confidence score is calculated for each pixel. The comprehensive confidence score is obtained by jointly calculating the pixel value of the pixel-level defect prediction image and the corresponding uncertainty metric. For each pixel, if the overall confidence score is higher than the preset score threshold, the pixel is assigned the value 1; otherwise, the pixel is assigned the value 0, resulting in a binary image. Connectivity analysis is performed on the binary image. Each connected component is determined using a four-adjacency method. The bounding rectangle of each connected component is extracted based on its spatial coverage. The center coordinates of the bounding rectangle of each connected component are then used as the location coordinates of the defect.
Citation Information
Patent Citations
Radiographic detection image defect identification method based on improved SwinT model
CN119887660A
SAR (Synthetic Aperture Radar) image cross-angle generation method and system based on guidance of physical scattering model
CN120580310A