Lung CT image segmentation method based on frequency spectrum Transform and ternary attention mechanism
By adopting a spectrum Transformer and ternary attention method in lung CT image segmentation, and using the parallel hybrid fusion module to extract image features, the problems of noise interference, insufficient data, category imbalance and low interpretability in lung CT image segmentation are solved, and higher segmentation accuracy and better model understanding ability are achieved.
Patent Information
- Application Number
- CN202510118148.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
In lung CT image segmentation, there are problems such as noise interference, lack of unified large-scale annotated data sets, image category imbalance, and low interpretability of deep learning models.
The lung CT image segmentation method based on spectrum Transformer and ternary attention is adopted. The parallel hybrid fusion module PHFM is combined with the spectrum Transformer module and the ternary attention module to extract global and local context information to enhance segmentation accuracy.
The accuracy of lung CT image segmentation is improved, especially in the recognition of lesion boundaries, which enhances the model's understanding of image features, improves the problem of category imbalance, and improves the interpretability of the model.
Smart Images

Figure CN120047464A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image segmentation, and relates to a method for segmenting lung CT medical images based on a spectral Transformer and a triple attention mechanism. Background Art
[0002] In the field of medical imaging, computed tomography (CT) technology has been widely popularized and applied. Especially in the diagnosis of lung diseases, CT scans provide high-resolution three-dimensional images, which can show the morphological changes of lung tissues and interstitial structures, and are beneficial to the discovery and localization of lung lesions. The segmentation of lung CT images can clearly distinguish normal tissues and diseased tissues in the lungs, and thus can more accurately judge the location and scope of the lesions. Therefore, the segmentation of lung CT images is an important research work, which has important value in improving the accuracy of diagnosis, assisting practitioners in making clinical decisions, reducing the workload of doctors, and promoting more in-depth research on diseases. First of all, the segmentation of lung CT images can improve the diagnostic accuracy. By effectively segmenting and extracting the lung parenchyma from lung CT images, it can better help doctors locate and analyze the lesion sites, diagnose lung diseases, and thus adopt more targeted treatment plans. In addition, the segmentation of lung CT images can effectively reduce the burden on doctors. In clinical practice, doctors need to read and interpret a large number of medical images, which is a time-consuming and highly professional task. By segmenting lung CT images, this process can be automated, enabling doctors to analyze and interpret images faster, thus saving valuable time and energy. Finally, the segmentation of lung CT images can promote the development of disease research. By segmenting lung CT images, researchers can better understand the development and progression of lung diseases, thus promoting the research on related diseases. For example, by segmenting the lung CT images of lung cancer patients, researchers can more accurately evaluate the size and spread of lung cancer, and thus better understand the biological characteristics and development trends of lung cancer.
[0003] The rapid development of deep learning technology has achieved remarkable results in the field of lung CT image segmentation, continuously promoting the improvement of segmentation accuracy. However, many problems still exist during the segmentation process, such as:
[0004] (1) Noise interference. When performing medical image cutting tasks, medical images are often affected by environmental factors, imaging instruments, or marking operation noises during acquisition or transmission.
[0005] (2) Lack of a unified large-scale labeled dataset, and the annotation quality of different datasets varies greatly. Due to many factors such as privacy protection, although there are many CT images of pneumonia, the vast majority are private data from healthcare providers, and public CT image data is relatively scarce. More datasets with challenging and different types of images are needed. The available image data of pneumonia patients is incomplete, noisy, blurred, and in some cases, the markings are inaccurate. It is very difficult to train a deep learning architecture with such insufficient medical datasets, and various problems (such as sparsity, missing values) must be solved. There is no standard and unified dataset for the use of datasets. In some works, some data augmentation methods have been proposed for this problem, such as semi-supervised methods to generate pseudo-labels and using transfer learning to learn the features of other diseases.
[0006] (3) Image class imbalance. The boundary between the infected area and the healthy area is relatively blurred, and the sizes of the lesions vary. In many cases, the infected area only accounts for a small part of the entire lung image, and this small part is the information that the segmentation task focuses on, resulting in a large difference in the pixel ratio between the ROI (Region of Infection) area and the surrounding tissue area, causing the problem of class imbalance. Therefore, there is a mismatch between positive and negative samples.
[0007] (4) The interpretability of applying deep learning to CT image segmentation is not high. Deep learning has performed well in many challenging benchmark tests. However, they still face some unsolved problems. A common problem with deep learning methods is the lack of transparency and interpretability, which may pose a serious threat and cause some limitations in clinical applications. Summary of the Invention
[0008] The first object of the present invention is to provide a method for segmenting lung CT medical images based on spectral Transformer and ternary attention, which is used to accurately segment the lesion part in medical images.
[0009] The specific implementation steps of the present invention are as follows:
[0010] Step 1: Preprocess and enhance the lung CT image data:
[0011] Collect a lung CT image dataset, generate two-dimensional axial slices and segmentation images using the original three-dimensional CT scan data and annotation information, and save them as pictures in a specific pixel format (such as png format with 512×512 pixels). Randomly divide the pictures into five equal parts, select one part as the validation set each time, and the rest as the training set. Use the five-fold cross-validation method to train the model and test it on the validation set, and calculate the average error.
[0012] Step 2: Construct a parallel hybrid fusion segmentation model PHFSeg:
[0013] The overall network architecture for lung CT image segmentation based on spectral Transformer and ternary attention consists of an encoder and a decoder. The encoder learns the multi-scale representation of the input image, and the decoder aggregates the representations at different levels of the encoder to predict the lung infection area.
[0014] In the encoder:
[0015] The encoder is composed of two parallel paths and nested skip paths, using the parallel hybrid fusion module PHFM as the basic module. The input is a lung CT slice (the grayscale image is replicated three times to become three channels), and it is downsampled four times to obtain images at different scales and processed in stages. In the first stage, the left side is the downsampling block DB, and the right side is the convolutional block CB. In other stages, PHFM undertakes the encoding. PHFM is composed of the spectral Transformer module STB and the ternary attention module TAM.
[0016] Design the spectral-based Transformer module STB, which combines the fast Fourier transform and Transformer. It captures different frequency components of the image through the spectral layer (including the FFT layer, weighted gating, and IFFT layer) to understand the local frequency. The weighted gating measures the importance of information with learnable weight parameters, determines the weights of the image lines and edges, and the learnable weight parameters are learned by backpropagation. After the spectral layer, there is a standard Transformer layer, including layer normalization, multi-head self-attention MHSA, and multi-layer perceptron MLP. Layer normalization helps to keep the gradient stable. MHSA maps the input to multiple feature subspaces for parallel processing to capture long-range dependencies. MLP performs non-linear transformation on the input data to help the model learn complex patterns.
[0017] In the spectral-based Transformer module STB, its operation process is that the original input is first layer-normalized, then after being transformed by the spectral layer, weighted gated, and inverse-transformed, it is added to the original data elements, and then processed by the Transformer layer through multiple layer normalizations, multi-head self-attention, and multi-layer perceptron and weighted to obtain the final output. The mathematical formula modeling can be expressed as:
[0018] O spe = X + ifft(fft(LN(X)) * W c )
[0019] O att = O spe + Attention(LN(O spe ))
[0020]
[0021] where fft(·) and ifft(·) represent the fast Fourier transform and its inverse transform, and W c represents the weight of the weighted gating, LN(·) represents layer normalization, Attention(·) represents multi-head attention, and MLP(·) represents a multi-layer perceptron. O spe and O att represent the output of the spectral layer and the output of the Transformer layer, respectively.
[0022] The triple attention module TAM is introduced. It adopts a unique three-way structure to capture cross-dimensional interaction and calculate attention weights. The input tensor establishes dependencies between different dimensions through rotation and residual connection, encoding channel and spatial information, and improving the model's ability to understand image features with low-cost calculations. The three parallel branches respectively capture the interactions between the channel and height, the channel and width, and the height and width dimensions, and the output is averaged and fused to act on the input data.
[0023] The triple attention module TAM consists of three parallel branches, each branch focusing on capturing the interaction between different dimensions. Specifically, two branches focus on the interaction between the channel dimension (C) and the spatial dimension (H or W), and the last branch is dedicated to establishing spatial attention. The outputs of these three branches are fused by taking the average and jointly act on the input data. For a given input tensor it is first passed to each of the three branches in the triple attention module. For the first branch, it focuses on establishing the interaction between the height dimension and the channel dimension. To achieve this, the input tensor X is rotated counterclockwise by 90° along the H axis. This operation rearranges the information of the original height dimension (H) to the channel dimension (C), and the information of the channel dimension is distributed on the height dimension. In this way, the first branch can specifically capture and model the mutual relationship between the channel dimension and the height dimension. This rotated tensor is denoted as with a shape of (W×H×C), and then is passed through the Z-Pool layer and its shape is transformed into (2×H×C) and then through a standard convolutional layer with a convolution kernel size of k×k, and then through a batch normalization layer. After this layer provides an intermediate output of dimension (1×H×C), this output passes through the Sigmoid activation function to generate attention weights, which finally act on the original tensor by rotating 90° clockwise along H and remaining the same shape as the original input X. Similarly, the second branch captures the interaction between the channel C and the width W, and the third branch establishes the interaction between the spatial dimensions H and W. By capturing this cross-dimensional interaction, the triple attention mechanism can effectively improve the model's ability to understand image features.
[0024] The Z-Pool layer effectively reduces the first dimension (usually the outermost dimension) of the input tensor to two by combining average pooling and max pooling features. This operation reduces the data depth without sacrificing the richness of the tensor representation, making subsequent processing more efficient. The specific mathematical formula is as follows:
[0025] Z-Pool(x) = [MaxPool 0d (x), AvgPool 0d (x)]
[0026] where 0d represents the 0th dimension (i.e., the outermost dimension) where max pooling and max average pooling occur. For example, a tensor of shape (C×H×W) outputs a tensor of size (2×H×W) after passing through the Z-Pool layer.
[0027] Similarly, in the second branch, the input X is rotated 90° counterclockwise along the W axis. The rotated tensor has a shape of (H×C×W), passes through the Z-Pool layer and outputs a of (2×C×W). Through a standard convolutional layer with a kernel size of k×k and a batch normalization layer, and then through a Sigmoid activation layer to generate attention weights and apply them to Finally, it is rotated 90° clockwise along W and remains the same as the original input shape X.
[0028] For the last branch, the number of channels of the input tensor X is reduced to two by the Z-Pool layer, and its shape is (2×H×W) Similarly, it also passes through a standard convolutional layer with a kernel size of k×k and a batch normalization layer, and through a Sigmoid activation layer to generate (1×H×W) attention weights and apply them to the input X tensor. Finally, the three outputs are aggregated, summed and averaged to output the final result of (C×H×W). To sum up, from the input tensor to the output result after triple attention The expression can be represented as:
[0029]
[0030] where σ represents the sigmoid activation function, and ψ 1 , ψ 2 and ψ 3 represent standard convolutional layers with a kernel size of k×k in the three branches.
[0031] In the decoder:
[0032] The upsampling process uses the PReLU activation function, upsamples the feature map proportionally in a bilinear interpolation manner, then uses the Feature Fusion Module (FFM) to fuse the upsampled feature maps of different scales with the feature maps at the corresponding encoder stages, and finally obtains the final prediction result through 1×1 convolution and batch normalization.
[0033] Step 3: Set the training strategy and loss function, and train the model;
[0034] The combination of weighted binary cross-entropy and Dice loss is used as the final loss function. The weighted binary cross-entropy is used to handle class-imbalanced data, and the weighted Dice loss function calculates the image similarity. The Adam optimizer is used and the decay weight is set. The Poly learning strategy is adopted to dynamically adjust the learning rate. After training for a certain number of iterations, the model performance is evaluated on the validation set, and the model parameters are adjusted according to the evaluation results. Finally, the generalization ability of the model is tested on the test set.
[0035] Step 4: Verify the trained network model:
[0036] The already segmented validation set is input into the trained segmentation model PHFSeg. After segmentation by the model, the diseased parts in the lung CT images are segmented out. The segmented images are compared and evaluated with the diseased areas judged by experts to verify the network model. The evaluation metrics include Intersection over Union (IoU), mean Intersection over Union (mIoU), Dice Similarity Coefficient (DSC), Sensitivity (SEN), Specificity (SPE), Precision (PRE), etc. The similarity between the prediction result and the ground truth label is quantified through these metrics to judge the model performance.
[0037] Step 5: Input the lung CT image to be segmented into the trained PHFSeg model, and the model outputs the segmentation result to obtain the prediction of the lung infection tissue area.
[0038] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects:
[0039] The present invention designs a parallel hybrid fusion module that fuses and calculates the spectral Transformer and triplet attention in a parallel manner, extracting global and local context information from the perspectives of sequence features and cross-dimensional features respectively. To better fuse the feature information calculated by the modules of the two parallel channels, it fuses the features of the two parallel channels through a multi-layer fusion method and returns them for the next calculation. Therefore, the present invention can extract global and local context information, thereby enhancing the segmentation accuracy, especially the boundary recognition of lung lesions. Among them, the encoder of the spectral Transformer module combines the long-range dependence ability of the Ttransformer and the ability to analyze spectral-based image information in the Fourier domain. The spectral-based Transformer module understands local frequencies by capturing different frequency components of the image, and uses the Transformer to capture relevant fine-grained features in the structure, alleviating the problem of loss of detail features caused by blurred segmentation boundaries or low contrast.
[0040] The present invention effectively combines the Transformer and the CNN, combines the function of the Transformer to capture global information through remote dependence with the function of the CNN to capture more detailed local information, enhances the functionality and flexibility of the traditional encoder-decoder architecture, and applies it to the field of medical image segmentation, realizing the automatic segmentation of the lesion part from the medical image. Brief Description of the Drawings
[0041] Figure 1 It is the structure diagram of the parallel hybrid fusion segmentation model PHFSeg of the present invention;
[0042] Figure 2 It is the module structure of the spectral Transformer module STB (Spectrum Transformer Block);
[0043] Figure 3 It is the structure diagram of the triplet attention module TAM (Triplet Attention);
[0044] Figure 4 It is the structure diagram of the parallel hybrid fusion module;
[0045] Figure 5 It is the medical image for testing;
[0046] Figure 6 It is Figure 5 The image after segmenting the lesion. Detailed Embodiment
[0047] The following further describes the present invention with reference to the accompanying drawings.
[0048] A lung CT image segmentation method based on spectral Transformer and ternary attention, specifically including the following steps:
[0049] Step 1: Preprocess and enhance the data of lung CT images:
[0050] In this example, two datasets, COVID-19-CT-Seg and MosMedData, are used. First, the data is preprocessed. Using the original three-dimensional CT scan data and annotation information, the corresponding two-dimensional axial slices and their segmentation images are generated and saved as png format images with a size of 512×512 pixels. Secondly, all the images are randomly divided into five equal parts, and each time one part is selected as the validation set, and the remaining four parts are used as the training set. Then, the images are rotated, flipped, translated, the brightness is changed, and multi-scale adjustments with different ratios (0.75, 1, 1.25) are performed to improve the generalization of the model.
[0051] Step 2: Construct the parallel hybrid fusion segmentation model PHFSeg:
[0052] The overall network architecture of lung CT image segmentation based on spectral Transformer and ternary attention consists of an encoder and a decoder. The encoder learns the multi-scale representation of the input image, and the decoder aggregates the representations at different levels of the encoder to predict the lung infection area.
[0053] As Figure 1 shown, the encoder of PHFSeg is composed of a series of nested skip paths on two parallel paths, and the Parallel Hybrid Fusion Module (PHFM) is used as the basic module among them.
[0054] In the encoder:
[0055] In the parallel structure of the encoder, the internal structures of the first stage and the other several stages are different. At the first stage (i.e., i = 1), the basic module on the left consists of a downsampling block (DB), and the basic module on the right consists of a convolutional block (CB). At the other several stages (i.e., i = 2, 3, 4), the parallel hybrid fusion module (PHFM) undertakes the encoding process. The parallel hybrid fusion module is composed of the spectrum-based Transformer module (STB) and the triple attention module (TAM) mentioned in the present invention. The parallel hybrid fusion module is not used in the first stage. Through experimental results, it is found that the designs of the downsampling block and the convolutional block in the first stage can better obtain global context information to avoid the loss of global information. The parallel hybrid fusion module uses a parallel method to simultaneously operate the spectrum Transformer module and the triple attention module, extracting global and local context information from the perspectives of sequence features and cross-dimensional features. The connection above and below each stage is completed through channel split. In the skip connection part of each stage, 1×1 convolution and batch normalization are performed to accelerate training and prevent overfitting.
[0056] As Figure 2 shown, the spectrum-based Transformer module (Spectrum Transformer Block, STB) designed by the present invention.
[0057] This module captures different frequency components of the image through the spectrum layer to understand local frequencies. The spectrum layer includes an FFT layer, a weighted gate, and an inverse fast Fourier transform (IFFT) layer. First, the FFT is used to convert the image from the physical space to the spectrum space. Second, the information flow is controlled through the weighted gate, and the learnable weight parameters are used to measure the importance of each frequency component to capture the lines and edges of the image. Finally, the IFFT is used to convert the spectrum space back to the physical space. After the spectrum layer is the standard Transformer layer, including layer normalization, multiple multi-head self-attention (MHSA), and a multi-layer perceptron (MLP). Layer normalization helps to keep the gradient stable and prevent gradient vanishing or explosion. MHSA maps the input to multiple feature subspaces, processes the data in parallel through multiple independent attention heads to capture long-range dependencies. The MLP is used for feature extraction and non-linear transformation to help the model learn complex patterns in the input data.
[0058] Its operation process is that the original input is first layer-normalized, added to the original data elements after being converted by the spectrum layer, weighted gated, and inverse-transformed, and then processed and weighted by the Transformer layer through multiple layer normalizations, multi-head self-attention, and multi-layer perceptrons to obtain the final output. The mathematical formula modeling can be expressed as:
[0059] O spe= X + ifft(fft(LN(X)) * W c )
[0060] O att = O spe + Attention(LN(O spe ))
[0061]
[0062] where fft(·) and ifft(·) represent the fast Fourier transform and its inverse transform respectively, W c represents the weight of the weighted gate, LN(·) represents layer normalization, Attention(·) represents multi-head attention, and MLP(·) represents a multi-layer perceptron. O spe and O att represent the output of the spectral layer and the output of the Transformer layer respectively.
[0063] As Figure 3 shown, the ternary attention structure designed by the present invention consists of three branches, and each branch focuses on capturing the interaction between different dimensions.
[0064] Two branches respectively focus on the interaction between the channel dimension (C) and the spatial dimension (H or W), and the last branch is dedicated to establishing spatial attention. The outputs of these three branches are fused by taking the average and act on the input data together.
[0065] For the input tensor the processing process of each branch is as follows:
[0066] The first branch rotates the input tensor X counterclockwise by 90° along the H axis to obtain a tensor with the shape of (W × H × C). After being processed by the Z-Pool layer, it is transformed into (with the shape of (2 × H × C)), and then the attention weight is generated through a standard convolutional layer and a batch normalization layer and applied to the original tensor after rotating it clockwise by 90° to keep the same shape as the original input.
[0067] The second branch rotates the input tensor X counterclockwise by 90° along the W axis to obtain a tensor (with the shape of (H × C × W)). Similarly, after being processed by the Z-Pool layer, it outputs (with the shape of (2 × C × W)), and the attention weight is generated and applied to X.
[0068] The third branch reduces the number of channels of the input tensor X to two through the Z-Pool layer and outputs (The shape is (2×H×W)). The attention weights are generated through the convolutional layer and the activation layer and applied to the input X.
[0069] Finally, the three outputs are aggregated, and the final result of (C×H×W) is output after summing and taking the average. In summary, from the input tensor to the output result after triple attention The expression can be represented as:
[0070]
[0071] where σ represents the sigmoid activation function, and ψ 1 , ψ 2 and ψ 3 represent standard convolutional layers with a convolutional kernel size of k×k in the three branches.
[0072] As Figure 4 shown, the parallel hybrid fusion module PHFM, which fuses the features of two parallel channels through multi-layer fusion and returns them for the next calculation. At the same time, residual connections are used between single channels. The multi-layer fusion structure retains the feature information obtained from the previous fusion calculation while calculating the feature information obtained from the single module of this layer.
[0073] For the path on the left side of the parallel structure, the present invention assumes that is the transformation function of the i-th stage and the k-th spectral Transformer block, where i ∈ {1, 2, 3, 4} and k ∈ {1, 2, …, M i}}, and its output result is denoted as Similarly, the present invention assumes that the transformation function of the i-th stage and the j-th triple attention block on the right path of the parallel structure is where i ∈ {1, 2, 3, 4} and j ∈ {1, 2, …, N i}}, and its output result is denoted as where C i represents the number of feature channels in the i-th stage. In the first stage, that is, the basic block when i = 1 is composed of a 3×3-based convolutional block and a downsampling convolution with 5×5 and stride = 2. Therefore, the first block in the first stage can be expressed by the formula:
[0074]
[0075] The first block in other stages can be expressed by the formula:
[0076]
[0077]
[0078] where \(i\in\{2,3,4\}\), represents a convolution operation with a convolution kernel of \(1\times1\). Split(\(\cdot\)) splits the feature map into two blocks along the channel dimension and passes them to and (\(i\in\{1,2,3,4\}\) here). The present invention uses a channel splitting module to fully integrate the features from two paths, because concatenation can avoid information loss in summation, thus performing feature fusion better than summation.
[0079] When extended to other blocks, it can be expressed as the formula:
[0080]
[0081] where \(i\in\{1,2,3,4\}\) and \(j\in\{2,\ldots,N\) i}\). Residual connections are used and the stride is 1. Additionally, when \(j - 1\leq M\) i , \(j'=j - 1\). When \(j'=M\) i , the left path can be derived as the formula:
[0082]
[0083] where \(i\in\{1,2,3,4\}\) and \(k\in\{2,\ldots,M\) i}\). Through Formula 3-3 and Formula 3-4, skip connections are established between the two paths of the encoder sub-network. This design helps the encoder perform multi-scale learning. Considering the balance among the number of network parameters, segmentation accuracy, and network performance, the present invention sets \(C\) i to \(\{8,24,32,64\}\), sets \(N\) i to \(\{3,4,9,7\}\), and sets \(M\) i to \(\{2,2,5,4\}\).
[0084] Decoder part:
[0085] During the upsampling process at each stage, the present invention uses the PReLU activation function to alleviate the vanishing gradient problem. PReLU introduces a learnable parameter \(\alpha\) to prevent the problem that neurons may die and never be activated. Since the input ratio of the top feature map of the encoder is \(1 / 16\) of the original input, the loss of fine details is disadvantageous for directly predicting the infected area. In the decoder sub-network, it is necessary to upsample and fuse the learned feature maps at each scale. The present invention adopts a Feature Fuse Module (FFM) for feature aggregation. At the same time, bilinear interpolation is used to upsample the feature map proportionally, and finally the final prediction result is obtained through \(1\times1\) convolution and batch normalization.
[0086] Step 3: Set the training strategy and loss function, and train the model;
[0087] Divide the preprocessed dataset into a training set, a test set, and a validation set. In the network model of the present invention, the backpropagation algorithm is used to update the weights and biases in the network; during the training iteration process, the loss function is used to update the parameters; in the selection of the loss function, a combination of weighted binary cross-entropy and Dice loss is adopted to train all networks;
[0088] The specific loss function is defined as:
[0089] Loss = Avg(L BCE + L Dice )
[0090] where Avg represents taking the average value, L BCE represents the weighted binary cross-entropy, and L Dice represents the weighted Dice loss function.
[0091] The definition of the weighted binary cross-entropy is:
[0092]
[0093] where l represents the label of two categories, 0 or 1; g i,j and p i,j respectively represent the true value and the predicted value corresponding to the coordinate (i, j) of the CT image with the size of H×W; ψ represents all the parameters in the model; Pr(·) represents the probability distribution of the prediction object; α i,j ∈[0, 1] represents the weight of each pixel. The weight α i,j can be defined as the formula:
[0094]
[0095] where A i,j represents the area around the pixel (i, j), and |·| represents the calculation of the absolute value.
[0096] The formula definition of the weighted Dice loss function is:
[0097]
[0098] Step 4: Verify the trained network model:
[0099] Input the already segmented validation set into the trained segmentation model PHFSeg. After the segmentation by the model, the lesion part in the medical image is segmented out, and the segmented image is compared and evaluated with the lesion area judged by experts;
[0100] Five widely adopted evaluation criteria are used to measure the performance of the PHFSeg model; the evaluation metrics are as follows:
[0101] The IoU defines the area of the intersection between the predicted segmentation map P and the ground truth map G, divided by the union area between the two maps, ranging from 0 to 1. The Mean Intersection over Union (mIoU) indicates the average of the IoU values for each category on this dataset.
[0102]
[0103] The Dice Similarity Coefficient (DSC) is a metric to measure the similarity between the predicted pulmonary infection (P) and the ground truth (G); TP, TN, FP, and FN represent true positive, true negative, false positive, and false negative respectively;
[0104]
[0105] Sensitivity: SEN represents the percentage of the lesion area that is correctly segmented;
[0106] Specificity: SPE represents the percentage of the non-lesion area that is correctly segmented;
[0107] Positive Predictive Value (Precision): PRE represents the accuracy of the lesion area segmentation,
[0108] Step Five: Input any medical image into the validated model and output the medical image of the segmented lesion:
[0109] Select a medical image for testing, as Figure 5 shown, this is the standard result of the segmented image, Figure 6 which is the result after the model segmentation. It can be seen that the PHFSeg model can accurately segment the lesion area, making it have an obvious boundary line with the surrounding normal tissue area. Such a result provides strong assistance for doctors to make diagnoses and treatments.
[0110] Compared with traditional medical image segmentation methods, the PHFSeg model has better performance, can more accurately segment the lesion area, and is also more applicable to different types of medical image datasets.
[0111] In the accompanying drawings of the disclosed embodiments of the present invention, only the structures related to the disclosed embodiments are involved. However, the above are only the preferred embodiments of the present invention, and it can be fully applied to various fields suitable for the present invention. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and the illustrated examples described herein.
Claims
1. A lung CT image segmentation method based on spectral Transformer and ternary attention, characterized in that: The method comprises the following steps: Step 1: preprocess and enhance lung CT image data; Step 2: Build the parallel hybrid fusion segmentation model PHFSeg: The parallel hybrid fusion segmentation model PHFSeg is composed of an encoder and a decoder; the encoder includes four stages, the first stage includes a nested skip-connected downsampling block DB and a convolution block CB, and a channel segmentation module; the second stage, the third stage, and the fourth stage are all composed of a parallel hybrid fusion module PHFM and a channel segmentation module; the parallel hybrid fusion module PHFM includes two parallel channels, the left channel includes multiple residual-connected spectrum Transformer modules STB, and the right channel includes multiple residual-connected ternary attention modules TAM; The decoder uses a bilinear difference method to upsample the feature map in proportion, and then uses a feature fusion module FFM to fuse the feature maps after upsampling at different scales with the feature maps of the corresponding encoder stage, and finally obtains the final prediction result through 1×1 convolution and batch normalization; Step 3: Set the training strategy and loss function to train the parallel hybrid fusion segmentation model PHFSeg; Step 4: Verify the trained parallel hybrid fusion segmentation model PHFSeg: Step 5: Input the lung CT image to be segmented into the trained parallel hybrid fusion segmentation model PHFSeg to obtain the prediction of the lung infection tissue area.
2. The lung CT image segmentation method based on spectral Transformer and ternary attention as claimed in claim 1, characterized in that: In the spectrum transformer module STB, the original input is first layer-normalized, and then added to the original input after spectrum layer conversion, weighted gating and inverse transformation, and then processed by multiple layers of transformer layer normalization, multi-head self-attention and multi-layer perceptron and weighted to obtain the final output; the formula of the spectrum transformer module STB is: O spe =X+ifft(fft(LN(X))*W c ) O att =P spe +Attention(LN(O spe )) Where fft(·) and ifft(·) represent the fast Fourier transform and its inverse transform, W c represents the weight of weighted gating, LN(·) represents layer normalization, Attention(·) represents multi-head attention, MLP(·) represents multi-layer perceptron, O spe and O att Represent the output of the spectral layer and the output of the Transformer layer respectively.
3. The lung CT image segmentation method based on spectral Transformer and ternary attention as claimed in claim 1 Cutting method, characterized in that: The ternary attention module TAM consists of three parallel branches. For a given input tensor It is input into three parallel branches respectively, and in the first branch and the second branch, it is respectively After transposition, Z-Pool, convolution, batch normalization, activation function, it is added to the original input and then transposed once more to obtain The outputs of the first and second branches are obtained, and then channel pooling, convolution, and batch normalization are performed in the third branch. After the activation function, the output of the third branch is obtained by adding it to the original input; the outputs of the three branches are aggregated. After summing and averaging, the final output is C×H×W; C is the channel dimension, H and W are the spatial dimensions Spend.
4. The lung CT image segmentation method based on spectral Transformer and ternary attention as claimed in claim 3 Cutting method, characterized in that: The formula for the Z-Pool layer is as follows: Z-Pool(x)=[MaxPool 0d (x),AvgPool 0d (x)] Among them, 0d represents the 0th dimension where the maximum pooling operation and maximum average pooling occur.
5. The lung CT image segmentation method based on spectral Transformer and ternary attention as claimed in claim 4 Cutting method, characterized in that: The triple attention module TAM takes the input tensor to pass The expression of the output result after ternary attention can be expressed as: in is the input after one transposition in the first branch, After Z-Pool operation is the input after one transposition in the second branch, After Z-Pool operation X is the original input, is X after the pooling operation, σ represents the sigmoid activation function, ψ1, ψ2 and ψ3 represent the three branches A standard convolutional layer with a kernel size of k×k. is the output result.
6. The lung CT image segmentation method based on spectral Transformer and ternary attention as claimed in claim 1 Cutting method, characterized in that: In the parallel hybrid fusion module PHFM, the input data flows into the left There are two parallel channels on the side and right sides. The spectrum information of the left channel is extracted through the spectrum transformer module STB The right channel extracts spatial and channel attention features through the ternary attention module TAM. Each STB and TAM performs feature transformation independently and fuses the features of the two paths through addition operation. Jump connection; The output of each layer is combined with the result of the previous fusion to form a layer-by-layer deep fusion. The layer fusion features are output from the last TAM; the output contains spectral information and attention enhancement features for the next calculation.
7. A lung CT image segmentation system implementing the method according to any one of claims 1 to 6, characterized in that: Includes the following modules: Data acquisition module, which acquires CT scan images of lung nodules and performs preprocessing; The image segmentation module uses the trained and verified parallel hybrid fusion segmentation model PHFSeg to segment the preprocessed lung nodule CT scan images.
8. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 6.
9. A computing device, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Lightweight image deblurring method based on improved Transform
CN118710538A
Mammary gland ultrasound image segmentation network based on diffusion probability model
CN118864493A