CT image processing method, device, CT imaging equipment and storage medium
By combining the FDK algorithm, USV-GAN network, and WAS-Mamba model, the problem of the lack of intelligent detection algorithms in traditional industrial CT imaging equipment has been solved, high-precision three-dimensional image reconstruction and target area segmentation have been achieved, and the ability to identify and quantitatively analyze tiny defects has been improved.
Patent Information
- Application Number
- CN202510672232.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-05-23
AI Technical Summary
Traditional industrial CT imaging equipment relies on manual experience for image interpretation, making it difficult to accurately locate and quantitatively analyze tiny defects and lacks intelligent detection algorithms.
A CT image processing method under sparse view is used, combined with the FDK algorithm, USV-GAN network and WAS-Mamba model, to perform three-dimensional image reconstruction and target area segmentation. Sparse view data is used to generate full view data, and image segmentation performance is improved through cross-channel window scanning and weighted state space model.
It achieves high-precision three-dimensional image reconstruction and target area segmentation, improves the ability to identify and quantitatively analyze tiny defects, and improves the accuracy and efficiency of automated detection.
Smart Images

Figure CN120182516B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of automated detection technology, and in particular to a CT image processing method, device, CT imaging equipment and storage medium. Background Art
[0002] Industrial CT imaging equipment is a non-destructive inspection device based on X-ray and computed tomography technologies. Its core function is to visualize the internal structure and defects of the object being inspected through high-resolution three-dimensional image reconstruction. Traditional defect identification relies primarily on manual visual interpretation of reconstructed images. This lacks intelligent detection algorithms tailored to the characteristics of industrial defects, making it difficult to accurately locate and quantitatively analyze small defects. Summary of the Invention
[0003] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention provides, on one hand, a CT image processing method, apparatus, CT imaging device and storage medium, which can reconstruct a three-dimensional image from sparse view image data;
[0004] Another aspect of the present invention provides a CT image processing method, apparatus, CT imaging device, and storage medium, which can segment a target area image from a reconstructed image.
[0005] In a first aspect, an embodiment of the present invention provides a CT image processing method, comprising:
[0006] Obtain the original projection data under sparse viewing angles;
[0007] Based on the FDK algorithm, a geometric weighting operation, a ramp filtering operation, and a smearing operation are performed on the original projection data to obtain first VVBP-Tensor data;
[0008] Sort the first VVBP-Tensor data by pixel value in the projection view direction to obtain second VVBP-Tensor data;
[0009] Based on the pre-trained USV-GAN network, the second VVBP-Tensor data is used to generate the third VVBP-Tensor data. The generator of the USV-GAN network adopts the UNet architecture combined with skip connections and stacking blocks to form a single-stage U-shaped network. The stacking block includes a multi-kernel depth-wise asymmetric convolution operation block, a non-activation block, and multiple layer normalization blocks.
[0010] An angular summation operation is performed on the third VVBP-Tensor data to obtain a three-dimensional reconstructed image.
[0011] According to some embodiments of the present invention, the pre-trained USV-GAN network generates third VVBP-Tensor data using the second VVBP-Tensor data, including:
[0012] Normalizing the second VVBP-Tensor data using the first layer normalization block to obtain first intermediate data;
[0013] performing feature extraction and fusion on the first intermediate data using the multi-kernel depthwise asymmetric convolution operation block to obtain second intermediate data, wherein the multi-kernel depthwise asymmetric convolution operation block is configured with three dilated convolution branches with different dilation rates;
[0014] Performing feature optimization on the second intermediate data using the non-activated block and then performing feature fusion and dimension adjustment using point convolution to obtain third intermediate data, wherein the non-activated block is configured with a first SimpleGate block, an SE block, and an SA block;
[0015] The third intermediate data and the second VVBP-Tensor data are element-wise added to obtain fourth intermediate data.
[0016] According to some embodiments of the present invention, performing element-wise addition on the third intermediate data and the second VVBP-Tensor data to obtain fourth intermediate data further includes:
[0017] The fourth intermediate data is normalized by the second layer normalization block and then sequentially processed by the first point convolution, the second SimpleGate block, and the second point convolution to obtain the third VVBP-Tensor data.
[0018] In a second aspect, an embodiment of the present invention provides a CT image processing method, comprising:
[0019] Obtain the original projection data under sparse viewing angles;
[0020] Based on the FDK algorithm and a pre-trained USV-GAN network, the original projection data is used to generate a three-dimensional reconstructed image. The generator of the USV-GAN network adopts a UNet architecture combined with skip connections and stacking blocks to form a single-stage U-shaped network. The stacking blocks include a multi-kernel depthwise asymmetric convolution operation block, a non-activation block, and multiple layer normalization blocks.
[0021] Inputting the three-dimensional reconstructed image into the WAS-Mamba model to segment the target area image;
[0022] Among them, the WAS-Mamba model includes multiple symmetrical encoder blocks and decoder blocks, each of which includes a window attention state space block for feature extraction. The window attention state space block adopts a residual network structure and is configured with a cross-channel window scanning layer and a weighted state space model.
[0023] According to some embodiments of the present invention, the cross-channel window scanning layer is used to:
[0024] The input feature map is divided into multiple windows of the same size, and the feature map is converted into four scan arrays by a window shift operation. The four scan arrays include the original feature map, the feature map rotated 180°, the feature map after the window is translated, and the feature map after translation and rotation;
[0025] The same scanning operation is performed on the four scanning arrays to obtain four scanning sequences, wherein the scanning operation includes: after completing the feature scanning in the first window, scanning the features in the same position window of the next channel.
[0026] According to some embodiments of the present invention, the weighted state-space model is used to:
[0027] Convolution is performed on the input feature map using a 1×1×1 convolution kernel to generate a spatial domain embedding sequence;
[0028] Performing Fourier transform on the input feature map and convolution processing using a 1×1×1 convolution kernel to generate a frequency domain embedding sequence;
[0029] Extracting features from the four scanning sequences using a state space block, and splicing them with the spatial domain embedding sequence and the frequency domain embedding sequence to obtain four spliced feature maps;
[0030] Use a 3×3 depth-wise separable convolution kernel to perform feature weighted screening on the four concatenated feature maps to obtain an intermediate feature map;
[0031] The intermediate feature map is subjected to dual-domain feature extraction using a 3×3 convolution kernel, and then the extracted features are subjected to dimensionality compression and an inverse window operation to obtain a WSSM feature image.
[0032] In a third aspect, an embodiment of the present invention provides a CT image processing device, comprising:
[0033] A first data acquisition module is used to acquire original projection data under a sparse viewing angle;
[0034] A first data processing module is configured to perform a geometric weighting operation, a slope filtering operation, and a smearing operation on the original projection data based on an FDK algorithm to obtain first VVBP-Tensor data;
[0035] a second data processing module, configured to sort the first VVBP-Tensor data by pixel value in a projection view direction to obtain second VVBP-Tensor data;
[0036] A third data processing module is used to generate third VVBP-Tensor data based on the pre-trained USV-GAN network using the second VVBP-Tensor data, wherein the generator of the USV-GAN network adopts a UNet architecture combined with skip connections and stacking blocks to form a single-stage U-shaped network, and the stacking blocks include a multi-kernel depthwise asymmetric convolution operation block, a non-activation block, and multiple layer normalization blocks;
[0037] The fourth data processing module is used to perform an angular summation operation on the third VVBP-Tensor data to obtain a three-dimensional reconstructed image.
[0038] In a fourth aspect, an embodiment of the present invention provides a CT image processing device, comprising:
[0039] A second data acquisition module acquires original projection data under a sparse viewing angle;
[0040] a fifth data processing module, configured to generate a three-dimensional reconstructed image using the raw projection data based on the FDK algorithm and a pre-trained USV-GAN network, wherein the generator of the USV-GAN network adopts a UNet architecture combined with skip connections and stacking blocks to form a single-stage U-shaped network, wherein the stacking blocks include a multi-kernel depthwise asymmetric convolution operation block, a no activation block, and multiple layer normalization blocks;
[0041] a sixth data processing module, configured to input the three-dimensional reconstructed image into a WAS-Mamba model to segment a target area image;
[0042] Among them, the WAS-Mamba model includes multiple symmetrical encoder blocks and decoder blocks, each of which includes a window attention state space block for feature extraction. The window attention state space block adopts a residual network structure and is configured with a cross-channel window scanning layer and a weighted state space model.
[0043] In a fifth aspect, an embodiment of the present invention provides a CT imaging device, comprising a processor and a memory, wherein the memory stores a computer program, and the processor is used to implement the above-mentioned CT image processing method when running the computer program.
[0044] In a sixth aspect, an embodiment of the present invention provides a storage medium, wherein the storage medium stores a computer program, and when the computer program is executed, the above-mentioned CT image processing method is implemented.
[0045] The embodiments of the present invention have at least the following beneficial effects:
[0046] On the one hand, the original projection data is processed based on the FDK algorithm, and the processed first VVBP-Tensor data is sorted. The sorted second VVBP-Tensor data contains structures similar to the target object and the structures in different slices are related. The USV-GAN network is then used to generate a third VVBP-Tensor data under full view based on the second VVBP-Tensor after processing through a multi-kernel depth-wise asymmetric convolution operation block, a no activation block, and multiple layer normalization blocks. The third VVBP-Tensor data is then summed in the angular direction, which can reconstruct a three-dimensional image from sparse view image data.
[0047] On the other hand, based on the FDK algorithm and the pre-trained USV-GAN network, the original projection data under sparse perspective is used to generate a three-dimensional reconstructed image, and then the WAS-Mamba model is used to segment the target area image. The WAS-Mamba model effectively retains local features by introducing a cross-channel window scanning layer and a weighted state space model and integrating them into a hierarchical U-shaped architecture, which is conducive to improving image segmentation performance.
[0048] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments with reference to the accompanying drawings, in which:
[0050] Figure 1 This is a flowchart of the steps of the CT image processing method according to Example 1 of the present invention;
[0051] Figure 2 This is a second step flow chart of the CT image processing method according to embodiment 1 of the present invention;
[0052] Figure 3 This is an architecture diagram of the USV-GAN network of Example 1 of the present invention;
[0053] Figure 4 for Figure 3 The architecture diagram of the USV-Block of the USV-GAN network is shown;
[0054] Figure 5 This is a flowchart of the steps of the CT image processing method according to Example 2 of the present invention;
[0055] Figure 6This is an architectural diagram of the WAS-Mamba model according to Example 2 of the present invention;
[0056] Figure 7 for Figure 6 The architecture diagram of WASBlock of the WAS-Mamba model is shown;
[0057] Figure 8 Schematic diagram of a window shifting scheme and scanning path for cross-channel window scanning according to embodiment 2 of the present invention;
[0058] Figure 9 A schematic diagram of a pixel-level display generated by a sequence of cross-channel window scanning according to embodiment 2 of the present invention;
[0059] Figure 10 for Figure 7 The architecture diagram of the weighted state space model of WASBlock is shown;
[0060] Figure 11 This is a principle block diagram of a CT image processing device according to embodiment 3 of the present invention;
[0061] Figure 12 This is a principle block diagram of a CT image processing device according to embodiment 4 of the present invention. DETAILED DESCRIPTION
[0062] The following describes embodiments of the present invention in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.
[0063] In the description of the present invention, "several" means one or more, "multiple" means more than two, "greater than," "less than," and "exceed" are understood to exclude the number itself, and "above," "below," and "within" are understood to include the number itself. The use of terms such as "first" and "second" is solely for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance, or implicitly specifying the number of the indicated technical features, or implicitly specifying the order of the indicated technical features.
[0064] Terminology / Translation:
[0065] CT: Computed Tomography, computed tomography;
[0066] FDK: Feldkamp-Davis-Kress, a classic method in cone-beam CT reconstruction;
[0067] VVBP-Tensor: View-by-View Backprojection Tensor, view-by-view backprojection tensor;
[0068] USV-GAN: Universal Sparse-view GAN, universal sparse view generative adversarial network;
[0069] UNet: A classic convolutional neural network (CNN) architecture for image segmentation.
[0070] Projections: projection data;
[0071] Log: logarithmic transformation;
[0072] Line integrals: curve integrals;
[0073] Ramp Filtering: Ramp filtering;
[0074] Smearing: smearing;
[0075] Sum: sum;
[0076] Reconstructed images: reconstructed images;
[0077] Before sorting: before sorting;
[0078] After sorting: after sorting;
[0079] Sorting: sorting operation;
[0080] Select longitudinal slices of VVBP-Tensor layer along angular direction: Select longitudinal slices of the VVBP-Tensor layer along the angular direction;
[0081] Sparse-view: sparse perspective;
[0082] Full-view: full view;
[0083] Concatenate: connect;
[0084] Input: input;
[0085] Output: output;
[0086] USV-Block: USV block;
[0087] DAC-Block: multi-kernel Depth-wise Asymmetric Convolution, multi-kernel depth-wise asymmetric convolution operation block;
[0088] dconv: dilated convolution, hole convolution;
[0089] LayerNorm, layer normalization;
[0090] AF-Block: Activation Free Block, no activation block;
[0091] SimpleGate: An improved gated linear unit, mainly used for image restoration tasks;
[0092] Conv: Convolution, convolution operation;
[0093] SE: Squeeze-and-Excitation, is an attention mechanism module used to enhance the performance of convolutional neural networks. It explicitly models the dependencies between channels and adaptively recalibrates channel feature responses, thereby improving the network's representation capabilities. The SE block consists of three core operations: Squeeze, Excitation, and Scale.
[0094] SA: Spatial Attention, a technique for enhancing the spatial information of feature maps in convolutional neural networks;
[0095] BN: Batch Normalization, batch normalization;
[0096] PatchGAN: is an improved GAN discriminator whose output is an N×N matrix, where each element x[i][j] represents a patch, corresponding to a receptive field of the image;
[0097] Reconstruction Loss: reconstruction loss;
[0098] Adversarial Loss: Adversarial loss;
[0099] Image: image;
[0100] Segmentation: Segment the image;
[0101] WAS: Windowed Attention State, window attention state;
[0102] SSM: State Space Model, state space model;
[0103] Mamba: is a new deep learning architecture based on state-space models;
[0104] WASBlock: Windowed Attention State Space Block;
[0105] WSSM: Weight State Space Model, weighted state space model;
[0106] CCWScan: Cross Channel Window Scan, cross channel window scan;
[0107] Patch Embedding: A layer in the encoder that splits the input data into fixed-size patches and converts them into embedding vectors.
[0108] Final Projection: A layer in the decoder that restores the size of features to match the size of the target segmentation features.
[0109] Patch Merging: A layer in the encoder that downsamples the input data.
[0110] Patch Expanding: A layer in the decoder used to upsample the input data;
[0111] Dconv: Depthwise Separable Convolution, depthwise separable convolution;
[0112] : Concatenation, connection;
[0113] : Element-wise Summation, element-level addition;
[0114] : Element-wise Production, element-wise multiplication;
[0115] VMamba: Vision Mamba, visual Mamba model;
[0116] SS2D: 2D-Selective-Scan, two-dimensional selective scanning;
[0117] Flatten: flatten;
[0118] Spatial Domain Embedding: Spatial domain embedding;
[0119] Frequency Domain Embedding: Frequency domain embedding;
[0120] Dual Domain Hybrid Embedding: Dual domain hybrid embedding;
[0121] Weighting of the Dual-Domain Auxiliary Branch: Weighting of the dual-domain auxiliary branch;
[0122] Embedded Branch: embedded branch;
[0123] Embedded Branch Fusion: embedded branch fusion;
[0124] Window Reversan&Reshape: Reverse window operation and reshape.
[0125] Example 1
[0126] Please refer to Figure 1 This embodiment discloses a CT image processing method, including S110 to S150. It should be noted that the numbering of the steps in this embodiment is only for the convenience of review and understanding, and does not limit the order of execution of the steps. The content of each step is detailed below:
[0127] S110, obtaining original projection data under a sparse viewing angle;
[0128] S120, performing a geometric weighting operation, a ramp filtering operation, and a smearing operation on the original projection data based on the FDK algorithm to obtain first VVBP-Tensor data;
[0129] S130 , sorting the first VVBP-Tensor data by pixel value in the projection view direction to obtain second VVBP-Tensor data;
[0130] S140, based on the pre-trained USV-GAN network, using the second VVBP-Tensor data to generate the third VVBP-Tensor data, the generator of the USV-GAN network adopts the UNet architecture combined with skip connections and stacking blocks to form a single-stage U-shaped network, the stacking block includes a multi-kernel depthwise asymmetric convolution operation block, a no activation block and multiple layer normalization blocks;
[0131] S150 , performing a summation operation in an angular direction on the third VVBP-Tensor data to obtain a three-dimensional reconstructed image.
[0132] Raw projection data can be acquired through CT imaging equipment. For industrial applications, the object being measured is placed on a turntable, the X-ray emitter and detector are activated, and the turntable rotates the object continuously. This allows for the acquisition of multiple projection images of the object at different angles, effectively acquiring the raw projection data. Currently, the angular positions of the projection images captured by CT imaging equipment are limited, meaning that the raw projection data is sparsely viewed.
[0133] Among the algorithms for CT image reconstruction, the FDK algorithm is a classic algorithm. The main steps of the FDK algorithm include geometric weighting, slope filtering, smearing, and summation. Please refer to Figure 2 The first and second rows of Figure 2 The first row shows that the original projection data (Projections) are pre-processed by logarithmic transformation (log) and line integrals, and then geometric weighting is performed (not shown in the figure). The weighted projection data is processed by ramp filtering (i.e., ramp filtering operation, Ramp Filtering), and the filtered data is smeared (also known as back-smearing or back-projection), where θ1, θ2....θ K Represents the smeared data at different angles, and sums the smeared data to synthesize the reconstructed images.
[0134] Since the images reconstructed by the FDK algorithm cannot meet the needs of practical applications, the algorithm needs to be further improved. The paper "VVBP-Tensor in the FBP Algorithm: Its Properties and Application in Low-Dose CT Reconstruction" proposed by XiTao et al. points out that the first VVBP-Tensor data obtained by the smearing operation can be sorted to obtain the second VVBP-Tensor data. The second VVBP-Tensor data contains structures similar to the target object and the structures in different slices are related. Figure 2 In the second and third rows, the first VVBP-Tensor data is sorted to obtain the second VVBP-Tensor data, and then the second VVBP-Tensor data is summed to obtain the reconstructed image.
[0135] However, the sorted second VVBP-Tensor data is still data from a sparse perspective, and the reconstructed image still cannot achieve a reconstructed image from a full perspective. Therefore, it is necessary to use the pre-trained USV-GAN data and the second VVBP-Tensor data to generate a third VVBP-Tensor data from a full perspective, and then sum the third VVBP-Tensor data to obtain a reconstructed image. Compared with the reconstructed image generated using the first VVBP-Tensor data or the second VVBP-Tensor data, the third VVBP-Tensor data is closer to full-view sampling, can fill in the missing angle information, and can achieve better reconstruction effects.
[0136] Step S140: Based on the pre-trained USV-GAN network, the second VVBP-Tensor data is used to generate third VVBP-Tensor data, including:
[0137] S141. Use the first layer normalization block to normalize the second VVBP-Tensor data to obtain first intermediate data;
[0138] S142. Perform feature extraction and fusion on the first intermediate data using a multi-kernel depthwise asymmetric convolution operation block to obtain second intermediate data, wherein the multi-kernel depthwise asymmetric convolution operation block is configured with three dilated convolution branches with different dilation rates.
[0139] S143. After optimizing the features of the second intermediate data using the non-activated block, feature fusion and dimension adjustment are performed using point convolution to obtain third intermediate data. The non-activated block is configured with the first SimpleGate block, the SE block, and the SA block.
[0140] S144 . Perform element-wise addition on the third intermediate data and the second VVBP-Tensor data to obtain fourth intermediate data.
[0141] After step S144, the following steps are also included:
[0142] S145. Use the second layer normalization block to normalize the fourth intermediate data, and then pass it through the first point convolution, the second SimpleGate block, and the second point convolution in sequence to obtain third VVBP-Tensor data.
[0143] The input of the USV-GAN network is the sorted second VVBP-Tensor data, and the output is the processed and restored third VVBP-Tensor data under the full perspective. The ultimate goal is to obtain the reconstructed image by summing these restored third VVBP-Tensor data along the angular direction.
[0144] For example, the USV-GAN model includes a generator and a discriminator. The generator takes the second VVBP-Tensor data of universal sparse view data as input. After a series of operations such as convolution, normalization, and attention mechanisms, it learns the internal structural features of a single layer of VVBP-Tensor and the relationships between different layers. This generates VVBP-Tensor data that can fill in the missing angle information, making the generated VVBP-Tensor data as close as possible to the VVBP-Tensor data sampled from the full view (the actual sampled data). The discriminator assists the generator training by evaluating the authenticity of smaller patches (patches) of the input image, prompting the generator to continuously optimize the output results to achieve better reconstruction results.
[0145] Please refer to Figure 3 , Figure 3 The generator of USV-GAN adopts UNet architecture combined with skip connection and stacking block to form a single-stage U-shaped network structure. The stacking block is a number of improved basic USV-Blocks (such as n 1. n 2... n 5). Please refer to Figure 4 ,Each USV-Block includes a multi-kernel depthwise asymmetric convolution operation block (DAC-Block), a no activation block (AF-Block) and multiple layer normalization blocks (LayerNorm).
[0146] Among them, the multi-kernel depthwise asymmetric convolution operation block (DAC-Block) uses dilated convolution (dconv) to expand the receptive field of the convolution kernel. The DAC-Block adopts a new three-branch architecture with different dilation rates for different branches (such as 5 for the second branch and 7 for the third branch). It simultaneously extracts local and contextual information, balancing the expansion of the receptive field with computational complexity. Finally, the feature maps generated by each branch are concatenated as the final output, helping the network learn the detailed internal structural features of the VVBP-Tensor.
[0147] Activation-free block (AF-Block): SimpleGate and a simplified channel attention mechanism are used to replace traditional activation functions, providing nonlinearity and capturing complex patterns and relationships. In addition, the spatial attention mechanism is combined to capture the spatial statistics of feature maps through pooling and convolution operations, giving higher weights to important areas, further improving model performance.
[0148] Layer Normalization (LayerNorm): Layer normalization is integrated into the basic block to alleviate the vanishing gradient problem by reducing the variance of activations in each layer, improving training stability and smoothness. Compared with other normalization techniques (such as BatchNormalization), it is less sensitive to feature scale;
[0149] Discriminator: PatchGAN is used to evaluate the authenticity of small patches of the input image rather than the entire image. This approach can more accurately locate real or fake regions, is robust to small changes in the input image, requires fewer parameters, and is more efficient to train.
[0150] Loss function: Two loss functions are used to optimize the model, namely reconstruction loss and adversarial loss. Reconstruction loss uses L1 distance to measure the generated image G ( x , y ) and the real image y The pixel-level difference is , E x,y,z To find the expectation of the joint distribution of input x, real image y and noise z, the function of the formula is to ensure that the generated image is close to the real image in pixel value, reduce image blur and make the generated image clearer.
[0151] The adversarial loss follows the standard non-saturated logistic GAN loss, which is formulated as
[0152]
[0153] During the training process, the discriminator D is trained to distinguish between real and generated images, while the generator G tries to generate images that the discriminator cannot distinguish. This adversarial learning method encourages the generator to continuously optimize the generated images and improve the image quality. The final optimization goal is
[0154]
[0155] The effects of these two loss functions on the model are comprehensively considered to achieve better cone-beam CT image reconstruction effects.
[0156] Example 2
[0157] Please refer to Figure 5 This embodiment provides a CT image processing method, including S210 to S230. It should be noted that the numbering of the steps in this embodiment is only for the convenience of review and understanding, and does not limit the order of execution of the steps. The content of each step is detailed below:
[0158] S210, obtaining original projection data under a sparse viewing angle;
[0159] S220, based on the FDK algorithm and the pre-trained USV-GAN network, generates 3D reconstructed images using raw projection data. The generator of the USV-GAN network adopts the UNet architecture combined with skip connections and stacking blocks to form a single-stage U-shaped network. The stacking blocks include multi-kernel depthwise asymmetric convolution operation blocks, no activation blocks, and multiple layer normalization blocks.
[0160] S230, inputting the 3D reconstructed image into the WAS-Mamba model to segment the target area image;
[0161] Among them, the WAS-Mamba model includes multiple symmetrical encoder blocks and decoder blocks. Both the encoder block and the decoder block include a window attention state space block for feature extraction. The window attention state space block adopts a residual network structure and is configured with a cross-channel window scanning layer and a weighted state space model.
[0162] Based on the FDK algorithm and a pre-trained USV-GAN network, a 3D reconstructed image is generated using raw projection data from a sparse perspective. The WAS-Mamba model is then used to segment the target area image. The WAS-Mamba model effectively preserves local features and improves image segmentation performance by introducing a cross-channel window scanning layer and a weighted state space model, integrating them into a layered U-shaped architecture. The specific methods of steps S210-S220 and the specific architecture of the USV-GAN network can be found in Example 1.
[0163] Please refer to Figure 6 The WAS-Mamba model adopts a U-shaped structure similar to U-Net, consisting of multiple symmetrical encoder and decoder modules. In the first encoder module, the Patch Embedding layer encodes the input image in blocks, and the Final Projection layer in the last decoder module restores the processed image to its original resolution. The remaining encoders gradually reduce the image resolution using Patch Merging layers, while the remaining decoders perform the reverse operation using Patch Expanding layers. Symmetrical skip connections are used between corresponding feature pyramids to reduce feature loss during resolution compression.
[0164] Each encoder module and decoder module contains a window attention state space block (WASBlock), which is the main feature extraction block and has a residual structure. Figure 7 , Figure 7The WASBlock architecture is shown. Layer normalization (Layer Norm) is first performed on the input feature map to generate two feature streams. The main feature stream uses linear layers and depthwise separable convolution (DConv) for initial feature extraction. Cross-channel window scanning (Cross-Channel Window Scan) then expands the feature map into a one-dimensional sequence, preserving local 3D features. A weighted state-space model (WSSM) is then used to weight the features of each sequence, completing feature extraction for the entire backbone network. Finally, linear layers are used to fuse the features and combine them with residual connections to form the output. Silu is the primary feature extraction block, with a residual structure.
[0165] Please refer to Figure 8 and Figure 9 , Figure 8 (a) is the window shift scheme; Figure 8 (b) is the cross-channel window scanning path, the yellow solid line marks the scanning path, the yellow dotted line marks the cross-channel path, and the blue dotted line marks the path to the next row of windows; Figure 9 Pixel-level display generated for the sequence. Cross-channel window scanning captures both local 3D and global features of the image by constructing the sequence using a cross-channel scanning path, achieving a balanced focus on local and global features during sequence conversion while minimizing the loss of original image features. Cross-channel window scanning preserves the two-dimensional features of the image when converting the image sequence in state space modeling, using this method to replace traditional scanning methods. Please refer to Figure 8 (a) divides the feature map into multiple windows of the same size (size is L×L), and the window shift step is L / 2; please refer to Figure 8 (b), after scanning the sequence in the first window, scan the features in the same position window of the next channel, and convert the feature map into four scan arrays through the window shift operation and scan to obtain four sequences. The four scan arrays include the original feature map (such as Figure 9 (a) of the ), the feature map after the translation window (as shown in Figure 9 (b)), the feature map after rotating 180° (as shown in Figure 9 (c)) and the feature map after translation and rotation by 180° (as shown in Figure 9 (d) is scanned to better capture the local 3D features and global features of the image and reduce the loss of original image features in sequence conversion, that is, the cross-channel window scanning layer is used to:
[0166] Please refer to Figure 9 and Figure 10 , the input feature map (such as Figure 10The input tag is shown as a mark) is divided into multiple windows of the same size, and the feature map is converted into four scan arrays by the window shift operation. The four scan arrays include the original feature map (such as Figure 9 or Figure 10 As shown in the mark (a)), the feature map after rotating 180° (as shown in Figure 9 or Figure 10 As shown in the mark (b)), the feature map after the translation window (as shown in Figure 9 or Figure 10 (c) in the figure) and the feature map after translation and rotation (as shown in Figure 9 or Figure 10 as indicated by the mark (d));
[0167] Perform the same scanning operation on the four scanning arrays to obtain four scanning sequences (such as Figure 10 The scanning operation includes: after completing the feature scanning in the first window, scanning the features in the same position window of the next channel.
[0168] The SS2D structure in traditional VMamba uses an equal processing strategy for all one-dimensional feature sequences, while the weighted state space model dynamically adjusts the weights according to the relative importance of the feature sequences. Figure 10 First, a hybrid embedding is generated through two branches. One branch uses 1×1×1 convolution to generate a spatial domain embedding of the input image, and the other branch uses Fourier transform (FFT) and 1×1×1 convolution to generate a frequency domain embedding. The four sequences after the cross-channel window scan are then concatenated with the embedded sequence after feature extraction using the state space model (SSM). Each concatenated embedded sequence is weighted and filtered using a 3×3 depthwise separable convolution (DConv). Finally, a 3×3 convolution is used to extract and compress the dual-domain features (frequency domain features and spatial domain features) to the original one-dimensional sequence size. The final feature image is generated through an inverse window operation, which can better identify key information and filter redundancy. Therefore, the weighted state space model is used to:
[0169] Use a 1×1×1 convolution kernel to convolve the input feature map to generate a spatial domain embedding sequence;
[0170] Perform Fourier transform on the input feature map and convolve it with a 1×1×1 convolution kernel to generate a frequency domain embedding sequence;
[0171] The state space block is used to extract features from the four scanning sequences and then concatenated with the spatial domain embedding sequence and the frequency domain embedding sequence to obtain four concatenated feature maps.
[0172] Use 3×3 depth-wise separable convolution kernel to perform feature weighted screening on the four spliced feature maps to obtain the intermediate feature map;
[0173] The 3×3 convolution kernel is used to extract dual-domain features from the intermediate feature map, and then the extracted features are dimensionally compressed and inverse windowed to obtain the WSSM feature image.
[0174] Among them, the loss function is based on the sum of the commonly used Dice loss and cross entropy loss, and the formula is:
[0175]
[0176] Where I represents the number of categories, V represents the number of voxels, and denote the true value and output probability of category i at voxel v, respectively.
[0177] Example 3
[0178] Please refer to Figure 11 This embodiment provides a CT image processing device, including:
[0179] A first data acquisition module 110 is used to acquire original projection data under a sparse viewing angle;
[0180] A first data processing module 120 is configured to perform a geometric weighting operation, a slope filtering operation, and a smearing operation on the original projection data based on the FDK algorithm to obtain first VVBP-Tensor data;
[0181] A second data processing module 130 is configured to sort the first VVBP-Tensor data by pixel value in a projection view direction to obtain second VVBP-Tensor data;
[0182] A third data processing module 140 is configured to generate third VVBP-Tensor data using the second VVBP-Tensor data based on a pre-trained USV-GAN network. The generator of the USV-GAN network adopts a UNet architecture combined with skip connections and stacking blocks to form a single-stage U-shaped network. The stacking blocks include a multi-kernel depthwise asymmetric convolution operation block, a non-activation block, and multiple layer normalization blocks.
[0183] The fourth data processing module 150 is configured to perform an angular summation operation on the third VVBP-Tensor data to obtain a three-dimensional reconstructed image.
[0184] The original projection data is processed based on the FDK algorithm, and the first VVBP-Tensor data obtained by the processing is sorted. The sorted second VVBP-Tensor data contains structures similar to the target object and the structures in different slices are related. The USV-GAN network is then used to generate third VVBP-Tensor data in full view based on the second VVBP-Tensor after processing through a multi-kernel depthwise asymmetric convolution operation block, a non-activation block, and multiple layer normalization blocks. The third VVBP-Tensor data is then summed in the angular direction to enable three-dimensional image reconstruction using sparse view image data. The inventive concept of this embodiment of the CT image processing device is the same as the inventive concept of the embodiment of the CT image processing method in the above-mentioned embodiment 1. Contents not covered in this embodiment of the CT image processing device can be referred to the embodiment of the CT image processing method in the above-mentioned embodiment 1 and will not be repeated here.
[0185] Example 4
[0186] Please refer to Figure 12 This embodiment provides a CT image processing device, including:
[0187] The second data acquisition module 210 acquires original projection data under a sparse viewing angle;
[0188] A fifth data processing module 220 is configured to generate a 3D reconstructed image using the original projection data based on the FDK algorithm and a pre-trained USV-GAN network. The generator of the USV-GAN network adopts a UNet architecture combined with skip connections and stacking blocks to form a single-stage U-shaped network. The stacking blocks include a multi-kernel depthwise asymmetric convolution operation block, a no activation block, and multiple layer normalization blocks.
[0189] The sixth data processing module 230 is used to input the 3D reconstructed image into the WAS-Mamba model to segment the target area image;
[0190] Among them, the WAS-Mamba model includes multiple symmetrical encoder blocks and decoder blocks. Both the encoder block and the decoder block include a window attention state space block for feature extraction. The window attention state space block adopts a residual network structure and is configured with a cross-channel window scanning layer and a weighted state space model.
[0191] Based on the FDK algorithm and a pre-trained USV-GAN network, a 3D reconstructed image is generated using raw projection data from a sparse perspective. The WAS-Mamba model is then used to segment the target area image. The WAS-Mamba model effectively preserves local features by introducing a cross-channel window scanning layer and a weighted state space model and integrating them into a layered U-shaped architecture, which helps improve image segmentation performance. The inventive concept of this CT image processing device embodiment is the same as the inventive concept of the CT image processing method embodiment of Example 2 above. For matters not covered in this CT image processing device embodiment, reference can be made to the CT image processing method embodiment of Example 2 above and will not be repeated here.
[0192] Example 5
[0193] This embodiment provides a CT imaging device including a processor and a memory. The memory stores a computer program. When the processor executes the computer program, it implements the above-described CT image processing method. The CT image processing method can be described with reference to the above-described embodiments 1 and 2 and will not be further described here.
[0194] On the one hand, the original projection data is processed based on the FDK algorithm, and the processed first VVBP-Tensor data is sorted. The sorted second VVBP-Tensor data contains structures similar to the target object and the structures in different slices are related. The USV-GAN network is then used to generate a third VVBP-Tensor data under full view based on the second VVBP-Tensor after processing through a multi-kernel depth-wise asymmetric convolution operation block, a no activation block, and multiple layer normalization blocks. The third VVBP-Tensor data is then summed in the angular direction, which can reconstruct a three-dimensional image from sparse view image data.
[0195] On the other hand, based on the FDK algorithm and the pre-trained USV-GAN network, the original projection data under sparse perspective is used to generate a three-dimensional reconstructed image, and then the WAS-Mamba model is used to segment the target area image. The WAS-Mamba model effectively retains local features by introducing a cross-channel window scanning layer and a weighted state space model and integrating them into a hierarchical U-shaped architecture, which is conducive to improving image segmentation performance.
[0196] Example 6
[0197] This embodiment provides a storage medium storing a computer program. When the computer program is executed, the above-mentioned CT image processing method is implemented. The CT image processing method can refer to the above-mentioned embodiments 1 and 2, and will not be described in detail here.
[0198] On the one hand, the original projection data is processed based on the FDK algorithm, and the processed first VVBP-Tensor data is sorted. The sorted second VVBP-Tensor data contains structures similar to the target object and the structures in different slices are related. The USV-GAN network is then used to generate a third VVBP-Tensor data under full view based on the second VVBP-Tensor after processing through a multi-kernel depth-wise asymmetric convolution operation block, a no activation block, and multiple layer normalization blocks. The third VVBP-Tensor data is then summed in the angular direction, which can reconstruct a three-dimensional image from sparse view image data.
[0199] On the other hand, based on the FDK algorithm and the pre-trained USV-GAN network, the original projection data under sparse perspective is used to generate a three-dimensional reconstructed image, and then the WAS-Mamba model is used to segment the target area image. The WAS-Mamba model effectively retains local features by introducing a cross-channel window scanning layer and a weighted state space model and integrating them into a hierarchical U-shaped architecture, which is conducive to improving image segmentation performance.
[0200] The embodiments of the present invention are described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Various changes can be made within the knowledge of ordinary technicians in the relevant technical field without departing from the scope of the present invention.
Claims
1. A CT image processing method, characterized in that: include: Obtain the original projection data under sparse viewing angles; Based on the FDK algorithm, a geometric weighting operation, a ramp filtering operation, and a smearing operation are performed on the original projection data to obtain first VVBP-Tensor data; Sort the first VVBP-Tensor data by pixel value in the projection view direction to obtain second VVBP-Tensor data; Based on the pre-trained USV-GAN network, the second VVBP-Tensor data is used to generate the third VVBP-Tensor data. The generator of the USV-GAN network adopts the UNet architecture combined with skip connections and stacking blocks to form a single-stage U-shaped network. The stacking block includes a multi-kernel depth-wise asymmetric convolution operation block, a non-activation block, and multiple layer normalization blocks. performing a summation operation in an angular direction on the third VVBP-Tensor data to obtain a three-dimensional reconstructed image; The pre-trained USV-GAN network generates third VVBP-Tensor data using the second VVBP-Tensor data, including: Normalizing the second VVBP-Tensor data using the first layer normalization block to obtain first intermediate data; performing feature extraction and fusion on the first intermediate data using the multi-kernel depthwise asymmetric convolution operation block to obtain second intermediate data, wherein the multi-kernel depthwise asymmetric convolution operation block is configured with three dilated convolution branches with different dilation rates; Performing feature optimization on the second intermediate data using the non-activated block and then performing feature fusion and dimension adjustment using point convolution to obtain third intermediate data, wherein the non-activated block is configured with a first SimpleGate block, an SE block, and an SA block; The third intermediate data and the second VVBP-Tensor data are element-wise added to obtain fourth intermediate data.
2. The CT image processing method according to claim 1, characterized in that: performing element-wise addition of the third intermediate data and the second VVBP-Tensor data to obtain fourth intermediate data, and then further comprising: The fourth intermediate data is normalized by the second layer normalization block and then sequentially processed by the first point convolution, the second SimpleGate block, and the second point convolution to obtain the third VVBP-Tensor data.
3. A CT image processing method, based on the CT image processing method according to claim 1 or 2, characterized in that: include: Obtain the original projection data under sparse viewing angles; Based on the FDK algorithm and a pre-trained USV-GAN network, the original projection data is used to generate a three-dimensional reconstructed image. The generator of the USV-GAN network adopts a UNet architecture combined with skip connections and stacking blocks to form a single-stage U-shaped network. The stacking blocks include a multi-kernel depthwise asymmetric convolution operation block, a non-activation block, and multiple layer normalization blocks. Inputting the three-dimensional reconstructed image into the WAS-Mamba model to segment the target area image; Among them, the WAS-Mamba model includes multiple symmetrical encoder blocks and decoder blocks, each of which includes a window attention state space block for feature extraction. The window attention state space block adopts a residual network structure and is configured with a cross-channel window scanning layer and a weighted state space model.
4. The CT image processing method according to claim 3, characterized in that: The cross-channel window scanning layer is used to: The input feature map is divided into multiple windows of the same size, and the feature map is converted into four scan arrays by a window shift operation. The four scan arrays include the original feature map, the feature map rotated 180°, the feature map after the window is translated, and the feature map after translation and rotation; The same scanning operation is performed on the four scanning arrays to obtain four scanning sequences, wherein the scanning operation includes: after completing the feature scanning in the first window, scanning the features in the same position window of the next channel.
5. The CT image processing method according to claim 4, characterized in that: The weighted state-space model is used to: Convolution is performed on the input feature map using a 1×1×1 convolution kernel to generate a spatial domain embedding sequence; Performing Fourier transform on the input feature map and convolution processing using a 1×1×1 convolution kernel to generate a frequency domain embedding sequence; Extracting features from the four scanning sequences using a state space block, and splicing them with the spatial domain embedding sequence and the frequency domain embedding sequence to obtain four spliced feature maps; Use a 3×3 depth-wise separable convolution kernel to perform feature weighted screening on the four concatenated feature maps to obtain an intermediate feature map; The intermediate feature map is subjected to dual-domain feature extraction using a 3×3 convolution kernel, and then the extracted features are subjected to dimensionality compression and inverse window operation to obtain a WSSM feature image.
6. A CT image processing device, characterized in that: include: A first data acquisition module is used to acquire original projection data under a sparse viewing angle; A first data processing module is configured to perform a geometric weighting operation, a slope filtering operation, and a smearing operation on the original projection data based on an FDK algorithm to obtain first VVBP-Tensor data; a second data processing module, configured to sort the first VVBP-Tensor data by pixel value in a projection view direction to obtain second VVBP-Tensor data; A third data processing module is configured to generate third VVBP-Tensor data based on a pre-trained USV-GAN network using the second VVBP-Tensor data, wherein the generator of the USV-GAN network adopts a UNet architecture combined with a skip connection and a stacking block to form a single-stage U-shaped network, wherein the stacking block includes a multi-kernel depth-wise asymmetric convolution operation block, a non-activation block, and a plurality of layer normalization blocks; the generation of the third VVBP-Tensor data based on the pre-trained USV-GAN network using the second VVBP-Tensor data includes: normalizing the second VVBP-Tensor data using the first layer normalization block to obtain first intermediate data; Performing feature extraction and fusion on the first intermediate data using the multi-kernel depthwise asymmetric convolution operation block to obtain second intermediate data, wherein the multi-kernel depthwise asymmetric convolution operation block is configured with three dilated convolution branches with different expansion rates; performing feature optimization on the second intermediate data using the non-activation block and then performing feature fusion and dimension adjustment using point convolution to obtain third intermediate data, wherein the non-activation block is configured with a first SimpleGate block, an SE block, and an SA block; performing element-wise addition on the third intermediate data and the second VVBP-Tensor data to obtain fourth intermediate data; The fourth data processing module is used to perform an angular summation operation on the third VVBP-Tensor data to obtain a three-dimensional reconstructed image.
7. A CT image processing device, based on the CT image processing device according to claim 6, characterized in that: include: A second data acquisition module acquires original projection data under a sparse viewing angle; a fifth data processing module, configured to generate a three-dimensional reconstructed image using the raw projection data based on the FDK algorithm and a pre-trained USV-GAN network, wherein the generator of the USV-GAN network adopts a UNet architecture combined with skip connections and stacking blocks to form a single-stage U-shaped network, wherein the stacking blocks include a multi-kernel depthwise asymmetric convolution operation block, a no activation block, and multiple layer normalization blocks; a sixth data processing module, configured to input the three-dimensional reconstructed image into a WAS-Mamba model to segment a target area image; Among them, the WAS-Mamba model includes multiple symmetrical encoder blocks and decoder blocks, each of which includes a window attention state space block for feature extraction. The window attention state space block adopts a residual network structure and is configured with a cross-channel window scanning layer and a weighted state space model.
8. A CT imaging device comprising a processor and a memory, wherein the memory stores a computer program, characterized in that: When the processor runs the computer program, it is used to implement the CT image processing method according to any one of claims 1 to 5.
9. A storage medium storing a computer program, wherein: When the computer program is executed, the CT image processing method according to any one of claims 1 to 5 is implemented.