Liver tumor segmentation method based on multidirectional spatial features
By using multi-directional spatial feature extraction and adaptive gating attention modules in liver tumor segmentation, the problem of high consumption of computing resources and difficulty in taking into account global semantic information and local detailed characteristics in the existing technology is solved, and the liver tumor segmentation effect with high precision, rapid response and efficient calculation is achieved.
Patent Information
- Application Number
- CN202510122987.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The prior art is difficult to ensure high segmentation accuracy, fast response ability and efficient computing performance in liver tumor segmentation, especially when processing 3D medical image data, and a single attention mechanism is difficult to take into account global semantic information and local detailed characteristics.
The liver tumor segmentation method based on multidirectional spatial characteristics is adopted, and the initial features are extracted and the adaptive gating attention module AGAM and the six-way Mamba module HoM are constructed, combined with the channel analytical entropy module CPEM, feature extraction, fusion and screening are performed, and the segmentation map of liver tumors is finally generated.
It realizes high-precision, fast response and efficient calculation of liver tumor segmentation, reduces the amount of model parameters and inference time, and improves the segmentation accuracy, and is suitable for complex medical image processing tasks.
Smart Images

Figure CN119991688A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image processing, and in particular to a liver tumor segmentation method based on multi-directional spatial features. Background Art
[0002] Liver tumors are abnormal cell groups in the liver, which seriously threaten the life and health of patients. Computerized tomography (CT) is an indispensable tool in modern clinical medicine, especially in the diagnosis of liver tumors. However, reading and analyzing CT images often requires doctors to have rich clinical experience and a lot of time and energy. For complex or small lesions, traditional manual reading methods may have the risk of missing, affecting the accuracy and timeliness of diagnosis.
[0003] Existing liver tumor segmentation methods are mainly divided into two categories: traditional machine learning and deep learning. Traditional machine learning methods are limited by the diversity of tumor morphology, size and location, as well as blurred boundaries, and are difficult to deal with complex medical image features, which limits their application. In recent years, deep learning has made significant progress in segmentation tasks. Its deep modeling capabilities for complex features can better capture complex semantic information, far surpassing traditional machine learning algorithms in many aspects.
[0004] In recent years, many studies on liver tumor segmentation have shown that the use of attention mechanisms can enhance the recognition of tumor edge features and micro-tumors, effectively alleviating the problems of blurred boundaries and inaccurate segmentation. However, the single attention mechanism still has shortcomings in modeling global semantic information and extracting local edge features. It is difficult to capture global features while fully taking into account the extraction of local detail features, which limits the further improvement of segmentation performance.
[0005] On the other hand, when processing 3D medical imaging data, directly using 3D convolution kernels can effectively extract spatial features, but it will consume a lot of computing resources. Although the Transformer-based method is good at capturing global information, its computational complexity increases quadratically with the feature dimension, and it also requires a lot of computing power.
[0006] Therefore, developing an automated liver tumor segmentation method that can ensure high segmentation accuracy while also having rapid response capabilities and efficient computing performance is of great clinical significance in helping doctors detect lesions earlier and formulate treatment plans in a timely manner. Summary of the invention
[0007] In order to solve the above technical problems, the present invention provides a liver tumor segmentation method based on multi-directional spatial features, which can provide users with a more accurate, efficient and highly adaptable liver tumor segmentation solution.
[0008] Specifically, the method comprises the following steps:
[0009] S1: Preprocess abdominal CT data to obtain NPY matrix data, and divide the data into training set, validation set and test set;
[0010] S2: Extract the initial features of NPY matrix data;
[0011] S3: Extract multi-directional spatial features from the initial features;
[0012] S4: Constructing the Adaptive Gated Attention Module AGAM e ) Extracting features from multi-directional spatial features to obtain local features;
[0013] S5: Construct a six-directional Mamba module HoM (Hexa-orientated Mamba) to extract multi-directional spatial features and obtain global features;
[0014] S6: local features and global features are fused and enhanced to obtain enhanced features;
[0015] S7: Construct a channel parsing entropy module CPEM (Channel Parsing Entropy Module) to screen the enhanced features and obtain the final features;
[0016] S8: Concatenate the final features at multiple scales to obtain the segmentation map of the liver tumor.
[0017] Preferably, S1 comprises the following steps:
[0018] S1.1: Collect the original abdominal CT data, perform preprocessing of standardization, cropping, and data enhancement on the data, and convert it into NPY matrix data f0;
[0019] S1.2: Divide the preprocessed data into training set, validation set and test set.
[0020] Preferably, S2 comprises the following steps:
[0021] A deep convolution is introduced to perform a convolution operation on the NPY matrix data f0 to obtain the initial feature f. The operation can be described as:
[0022] f = DepthwiseConv(f0, s, p)
[0023] Among them, DepthwiseConv(.,.,.) represents a depth convolution with a stride of s and a padding size of p.
[0024] Preferably, S3 comprises the following steps:
[0025] Construct a multi-directional spatial feature extractor MDFE (Multi Directional Feature Extractor). The operation can be described as:
[0026] A set of spatial separation convolutions are introduced to perform continuous convolution operations on the initial feature f along the horizontal, vertical and depth directions to obtain the multi-directional spatial feature f w , which is expressed as follows:
[0027] f h =Conv 3×1×1 (f)
[0028] f v =Conv 1×3×1 (f h )
[0029] f w =Conv 1×1×3 (f v )
[0030] In the above formula, f h represents the output of f after horizontal convolution, f v Represents f h The output after vertical convolution, f w represents f v The output after depthwise convolution represents multi-directional spatial features. 3×1×1 (.) represents horizontal convolution calculation, Conv 1×3×1 (.) represents the vertical convolution calculation, Conv 1×1×3 (.) represents the depth-wise convolution calculation.
[0031] Preferably, S4 comprises the following steps:
[0032] S4.1: Extract features from cross-sectional slices of multi-directional spatial features;
[0033] S4.2: Calculate attention weights;
[0034] S4.3: Perform local attention enhancement on multi-directional spatial features to obtain local features f AGAM .
[0035] Preferably, S4.1 comprises the following steps:
[0036] fl =Conv(f w )
[0037] In the above formula, fl represents the cross-sectional characteristics;
[0038] Preferably, S4.2 comprises the following steps:
[0039] The multi-directional spatial features are up-sampled and down-sampled, and the attention weights are generated through convolution.
[0040] The operation can be described as:
[0041] f down = Downsampled(f w )
[0042] f coe =Conv(f down )
[0043] f up =Upsampled(f coe )
[0044] α=Sigmoid(f up )
[0045] In the above formula, f down Represents f w The output after downsampling operation, f coe Represents f down The output after the convolution operation, f up Represents f coe The output after upsampling operation, α represents the attention weight, Downsampled(.) represents downsampling, which means reducing the spatial resolution of the input, Upsampled(.) represents upsampling, which means increasing the spatial resolution of the input, and Sigmoid(.) represents normalization;
[0046] Preferably, S4.3 comprises the following steps:
[0047] The attention weight α is used to adjust the multi-directional spatial features and fuse them with the cross-sectional features fl. This operation can be described as:
[0048] f AGAM =f l +αf w
[0049] In the above formula, αf w represents the use of attention weights to adjust multi-directional spatial features, f AGAM Represents the output of the adaptive gated attention module AGAM.
[0050] Preferably, S5 comprises the following steps:
[0051] S5.1: Construct Mamba long sequence modeling operation;
[0052] S5.2: Flatten the coronal, sagittal and transverse sequences of multi-directional spatial features into Mamba long sequence models.
[0053] Preferably, S5.1 comprises the following steps:
[0054] Given the feature dimension t, define f t ={x1, x2, ..., x t}, f′ t ={y′1,y′2,...,y′ t}, f′ t =Mamba(f t );
[0055] Among them, x t 2D flattened sequence representing features, f t Representative feature, f′ t Representative feature f t The output after Mamba long sequence modeling, y′ t represents f′ t 2D flattened sequence, Mamba(.) represents the Mamba long sequence modeling operation;
[0056] The Mamba long sequence modeling operation Mamba(.) is defined as follows:
[0057] x t After linear transformation and convolution operation, the expression is:
[0058] x t =Conv(Linear(x t ))
[0059] In the above formula, Linear(.) represents linear transformation;
[0060] Introducing parameter B t and C t , whose expression is:
[0061] B t =Linear B (x t )
[0062] C t =Linear C (x t )
[0063] In the above formula, Bt Represents according to x t The dynamically generated input transformation matrix, C t Represents according to x t Dynamically generated input transformation matrix, Linear B (.) stands for B t Linear transformation, Linear C (.) stands for C t Linear transformation of ;
[0064] Construct the state space model, which is expressed as:
[0065] h t =Ah t-1 +B t x t
[0066] y t =C t h t
[0067] In the above formula, h t represents the hidden state of the state space model, y t represents the predicted output of the state space model, and A represents the state transfer matrix;
[0068] Based on the state space model, the dynamic time step Δ is introduced t , whose expression is:
[0069] Δ t =Linear Δ (x t )
[0070] In the above formula, Linear Δ (.) represents a special case for Δ t Linear transformation of ;
[0071] Construct a time-varying state space model, which is expressed as:
[0072] A′=exp(Δ t A)
[0073] B′=(Δ t A) -1 (exp(Δ t A)-I)B t
[0074] In the above formula, A′, B′ represent the time step Δ t After discretization, I represents the diagonal matrix, and exp(.) represents the natural exponential function;
[0075] h′t =A′h′ t-1 +B′x t
[0076] y′ t =C t h′ t
[0077] In the above formula, h′ t represents the hidden state of the time-varying state space model, y′ t Represents the predicted output of the time-varying state-space model:
[0078] The multi-directional spatial feature f w After Mamba long sequence modeling operation, we can get f′ w , whose expression is:
[0079] f′ w =Mamba(f w )
[0080] In the above formula, f′ w Represents the multi-directional spatial feature f w Output after Mamba long sequence modeling;
[0081] Preferably, S5.2 comprises the following steps:
[0082] The multi-directional spatial features are flattened into 6 feature sequences, and Mamba long sequence modeling is performed on these 6 sequences respectively. Then, the modeled sequences are fused to achieve global attention enhancement of the multi-directional spatial features. This operation can be described as:
[0083] f Mamba =Mamba(f as )+Mamba(f ps )+Mamba(f sc )+Mamba(f ic )+Mamba(f la )+Mamba(f ra )
[0084] In the above formula, f Mamba represents the output of the six-way Mamba module HoM, f as Represents the multi-directional spatial feature f w Anterior flattening sequence of sagittal plane, f ps Represents the multi-directional spatial feature f w Posterior flattening sequence of sagittal plane, f sc Represents the multi-directional spatial feature f w The flattened sequence of the coronal plane, f ic Represents the multi-directional spatial feature fw The inferior flattened sequence of the coronal plane, f la Represents the multi-directional spatial feature f w The left-hand flattened sequence of the cross section, f ra Represents the multi-directional spatial feature f w Right-hand flattened sequence of cross sections.
[0085] Preferably, S6 comprises the following steps:
[0086] S6.1: Fuse the outputs of the adaptive gated attention module AGAM and the six-way Mamba module HoM to obtain the fused feature f merge ;
[0087] S6.2: Enhance the fused features to obtain enhanced features.
[0088] Preferably, S6.1 comprises the following steps:
[0089] The multi-directional spatial features enhanced by the six-way Mamba attention module HoM and the adaptive gated attention module AGAM are fused to obtain the fused feature f merge , the operation can be described as:
[0090] f merge =f AGAM +f Mamba
[0091] In the above formula, f merge represents fusion features;
[0092] Preferably, S6.2 comprises the following steps:
[0093] For the fusion feature f merge Perform layer normalization to balance the distribution of features, and then use a multi-layer perceptron to perform nonlinear transformation and feature mapping to further enhance the expressiveness of features and obtain the enhanced feature f s , the operation can be described as:
[0094] f s =MLP(LN(f merge ))+f merge
[0095] In the above formula, f s stands for enhanced features, MLP(.) stands for multi-layer perceptron, and LN(.) stands for layer normalization.
[0096] Preferably, S7 comprises the following steps:
[0097] S7.1: Calculate the channel selectivity coefficient of the enhanced feature;
[0098] S7.2: Calculate uncertainty estimates for enhanced characteristics;
[0099] S7.3: Construct channel parsing entropy module CPEM to screen and enhance features.
[0100] Preferably, S7.1 comprises the following steps:
[0101] The input enhanced features are globally averaged and pooled, and the enhanced features of each channel are mapped to a scalar, which is then input into the multi-layer perceptron. The importance of each channel is modeled through nonlinear transformation, and the dependencies between channels are further extracted. Finally, normalization is performed to map the important coefficients of each channel to the range of [0, 1] to generate channel selection coefficients. This operation can be described as:
[0102] W = Sigmoid(MLP(GAP(f s )))
[0103] In the above formula, W represents the channel selection coefficient, and GAP(.) represents the global average pooling;
[0104] Preferably, S7.2 comprises the following steps:
[0105] Given the current feature dimension C, calculate the channel average of the current dimension to obtain the normalized representation f of the single-scale feature sigle , the operation can be described as:
[0106]
[0107] In the above formula, represents the enhanced features of the i-th channel;
[0108] The information entropy formula is used to calculate the uncertainty estimation coefficient u of the feature. This operation can be described as:
[0109] u=-f s log(f s )
[0110] In the above formula, u is a constant, -f s log(f s ) represents information entropy, which measures the uncertainty of the feature. A larger entropy value indicates a higher uncertainty of the feature.
[0111] Preferably, S7.3 comprises the following steps:
[0112] Adjust each channel feature, integrate the channel importance weight W and uncertainty estimation coefficient u, and obtain the filtered features. This operation can be described as:
[0113] f out =Wfs +(1-u)f s
[0114] In the above formula, f out Represents the final feature.
[0115] Preferably, S8 comprises the following steps:
[0116] S8.1: Concatenate the final features of each scale to obtain the final feature map;
[0117] S8.2: Construct a segmentation head to convert the final feature map into a segmentation map;
[0118] S8.3: Use the training set, validation set, and test set to train, validate, optimize, and test the model.
[0119] Preferably, S8.2 comprises the following steps:
[0120] f class =Conv(f out )
[0121]
[0122] In the above formula, f class Represents f out The output after the convolution operation, f class [i, j, k, c] represents f class The pixel (i, j, k, c), i represents the horizontal coordinate, j represents the vertical coordinate, k represents the vertical coordinate, c represents the category of the pixel, P i,j,k,c represents the probability that the pixel (i, j, k, c) belongs to category c, argmax represents the output category label based on the maximum probability, and F[i, j, k] represents the value of the (i, j, k) pixel of the segmentation map.
[0123] The beneficial effects of the present invention are as follows: first, in the data preprocessing stage, by standardizing, cropping, and data enhancement operations on the collected liver tumor images, the data quality is improved, the interference of irrelevant information is reduced, and the data is reasonably divided into a training set, a validation set, and a test set, laying a good foundation for subsequent model training and testing. In the process of model construction, a spatial feature extractor MDFE is proposed to strengthen the spatial connection of features. Then, an adaptive gated attention module AGAM is proposed, which realizes the capture of local detail features while using a small number of parameters. Next, a six-way Mamba module HoM is proposed, which can well capture global features while using a small number of parameters. Subsequently, a channel parsing entropy module CPEM is proposed to remove invalid features. These innovative designs enable the model to maintain a high level of segmentation accuracy while significantly reducing the number of parameters and speeding up the inference speed, providing an efficient and reliable solution for liver tumor segmentation tasks, showing good clinical application potential. BRIEF DESCRIPTION OF THE DRAWINGS
[0124] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the prior art and the drawings required for use in the embodiments. The following drawings are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0125] Figure 1 It is a flowchart of a liver tumor segmentation method based on multi-directional spatial features of the present invention;
[0126] Figure 2 It is a model architecture diagram of a liver tumor segmentation method based on multi-directional spatial features of the present invention;
[0127] Figure 3 It is a schematic diagram of the structure of an adaptive gated attention module AGAM of a liver tumor segmentation method based on multi-directional spatial features of the present invention;
[0128] Figure 4 It is a schematic diagram of the six-directional Mamba module HoM structure of a liver tumor segmentation method based on multi-directional spatial features of the present invention;
[0129] Figure 5 It is a schematic diagram of the structure of a channel parsing entropy module CPEM of a liver tumor segmentation method based on multi-directional spatial features of the present invention; Specific implementation plan
[0130] In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all other embodiments obtained by ordinary technicians in this field without creative work based on the embodiments of the present invention are within the scope of protection of the present invention.
[0131] The embodiment of the present invention provides a liver tumor segmentation method based on multi-directional spatial features, which is used to achieve accurate segmentation of liver tumors, assist clinical diagnosis and treatment decisions, and improve the accuracy and efficiency of medical image processing.
[0132] Reference Figure 1 , the method comprises the following steps:
[0133] S1: Preprocess abdominal CT data to obtain NPY matrix data, and divide the data into training set, validation set and test set;
[0134] S2: Extract the initial features of NPY matrix data;
[0135] S3: Extract multi-directional spatial features from the initial features;
[0136] S4: Construct an adaptive gated attention module AGAM to extract multi-directional spatial features and obtain local features;
[0137] S5: Construct a six-way Mamba module HoM to extract multi-directional spatial features and obtain global features;
[0138] S6: local features and global features are fused and enhanced to obtain enhanced features;
[0139] S7: Construct a channel parsing entropy module CPEM to screen the enhanced features and obtain the final features;
[0140] S8: Concatenate the final features at multiple scales to obtain the segmentation map of the liver tumor.
[0141] Further, S1 comprises the following steps:
[0142] S1.1: Collect the original abdominal CT data, perform preprocessing of standardization, cropping, and data enhancement on the data, and convert it into NPY matrix data f0;
[0143] Furthermore, the step of collecting the original abdominal CT data, preprocessing the data by standardization, cropping, and data enhancement, and converting the data into NPY matrix data f0 specifically includes:
[0144] The collected data include public data sets and data sets provided by cooperative hospitals. Each case image completely covers the entire area of the liver. The CT data are cropped to remove irrelevant areas, resampled to 512×512×512 size, and the pixel value of each data is normalized to a uniform range (0 to 1), and then the data is converted into NPY matrix format f0;
[0145] S1.2: Divide the preprocessed data into training set, validation set and test set;
[0146] Furthermore, the step of dividing the preprocessed data into a training set, a validation set, and a test set specifically includes:
[0147] The data is divided into training set, validation set and test set in a ratio of 8:1:1.
[0148] Further, S2 comprises the following steps:
[0149] A 7×7×7 convolution kernel is introduced, and the NPY matrix data f0 is convolved with a stride of 2×2×2 and padded to 3×3×3 to obtain the initial feature f. The operation can be described as:
[0150] f = DepthwiseConv(f0, s, p)
[0151] Among them, DepthwiseConv(.,.,.) represents a depth convolution with a stride of s and a padding size of p.
[0152] Further, S3 includes the following steps:
[0153] Construct a multi-directional spatial feature extractor MDFE (Multi Directional Feature Extractor). The operation can be described as:
[0154] A set of spatial separation convolutions are introduced to perform continuous convolution operations on the initial feature f along the horizontal, vertical and depth directions to obtain the multi-directional spatial feature f w , which is expressed as follows:
[0155] f h =Conv 3×1×1 (f)
[0156] f v =Conv 1×3×1 (f h )
[0157] f w =Conv 1×1×3 (f v )
[0158] In the above formula, f h represents the output of f after horizontal convolution, f v Represents f h The output after vertical convolution, f w represents f v The output after depthwise convolution represents multi-directional spatial features. 3×1×1 (.) represents horizontal convolution calculation, Conv 1×3×1 (.) represents the vertical convolution calculation, Conv 1×1×3 (.) represents the depth-wise convolution calculation.
[0159] Further, refer to Figure 3 , S4 comprises the following steps:
[0160] S4.1: Extract features from cross-sectional slices of multi-directional spatial features;
[0161] S4.2: Calculate attention weights;
[0162] S4.3: Perform local attention enhancement on multi-directional spatial features to obtain local features f AGAM .
[0163] Further, S4.1 includes the following steps:
[0164] The convolution operation is performed using a convolution kernel of size 3×3×1.
[0165] f l =Conv(f w )
[0166] In the above formula, f l Represents cross-sectional characteristics;
[0167] Further, S4.2 includes the following steps:
[0168] The multi-directional spatial features are up-sampled and down-sampled, and the attention weights are generated through convolution.
[0169] The operation can be described as:
[0170] f down = Downsampled(f w )
[0171] f coe =Conv(f down )
[0172] f up =Upsampled(f coe )
[0173] α=Sigmoid(f up )
[0174] In the above formula, f down Represents f w The output after downsampling operation, f coe Represents f down The output after the convolution operation, f up Represents f coe The output after upsampling operation, α represents the attention weight, Downsampled(.) represents downsampling, which means reducing the spatial resolution of the input, Upsampled(.) represents upsampling, which means increasing the spatial resolution of the input, and Sigmoid(.) represents normalization;
[0175] Further, S4.3 includes the following steps:
[0176] The multi-directional spatial features are adjusted using the attention weights and fused with the cross-sectional features fl to achieve local attention enhancement of the multi-directional spatial features. This operation can be described as:
[0177] f AGAM =f l +αf w
[0178] In the above formula, αf w represents the use of attention weights to adjust multi-directional spatial features, f AGAM Represents the output of the adaptive gated attention module AGAM.
[0179] Further, refer to Figure 4 , S5 comprises the following steps:
[0180] S5.1: Construct Mamba long sequence modeling operation;
[0181] S5.2: Model the coronal, sagittal, and transverse flattened sequences and the retrograde Mamba long sequence of multi-directional spatial features.
[0182] Further, S5.1 includes the following steps:
[0183] Given the feature dimension t, define f t ={x1, x2, ..., x t}, f′ t ={y′1,y′2,...,y′ t}, f′ t =Mamba(f t );
[0184] Among them, x t2D flattened sequence representing features, f t Representative feature, f′ t Representative feature f t The output after Mamba long sequence modeling, y′ t represents f′ t 2D flattened sequence, Mamba(.) represents the Mamba long sequence modeling operation;
[0185] The Mamba long sequence modeling operation Mamba(.) is defined as follows:
[0186] x t After linear transformation and convolution operation, the expression is:
[0187] x t =Conv(Linear(x t ))
[0188] In the above formula, Linear(.) represents linear transformation;
[0189] Introducing parameter B t and C t , whose expression is:
[0190] B t =Linear B (x t )
[0191] C t =Linear C (x t )
[0192] In the above formula, B t Represents according to x t The dynamically generated input transformation matrix, C t Represents according to x t Dynamically generated input transformation matrix, Linear B (.) stands for B t Linear transformation, Linear C (.) stands for C t Linear transformation of ;
[0193] Construct the state space model, which is expressed as:
[0194] h t =Ah t-1 +B t x t
[0195] y t =C t h t
[0196] In the above formula, h t represents the hidden state of the state space model, y t represents the predicted output of the state space model, and A represents the state transfer matrix;
[0197] Based on the state space model, the dynamic time step Δ is introduced t , whose expression is:
[0198] Δ t =Linear Δ (x t )
[0199] In the above formula, Linear Δ (.) represents a special case for Δ t Linear transformation of ;
[0200] Construct a time-varying state space model, which is expressed as:
[0201] A′=exp(Δ t A)
[0202] B′=(Δ t A) -1 (exp(Δ t A)-I)B t
[0203] In the above formula, A′, B′ represent the time step Δ t After discretization, I represents the diagonal matrix, and exp(.) represents the natural exponential function;
[0204] h′ t =A′h′ t-1 +B′x t
[0205] y′ t =C t h t
[0206] In the above formula, h′ t represents the hidden state of the time-varying state space model, y′ t represents the predicted output of the time-varying state-space model;
[0207] The multi-directional spatial feature f w After Mamba long sequence modeling operation, we can get f′ w , whose expression is:
[0208] f w =Mamba(f w )
[0209] In the above formula, f′ w Represents the multi-directional spatial feature f w Output after Mamba long sequence modeling;
[0210] Further, S5.2 includes the following steps:
[0211] The multi-directional spatial features are flattened into 6 feature sequences, and Mamba long sequence modeling is performed on these 6 sequences respectively. Then, the modeled sequences are fused to achieve global attention enhancement of the multi-directional spatial features. This operation can be described as:
[0212] f Mamba =Mamba(f as )+Mamba(f ps )+Mamba(f sc )+Mamba(f ic )+Mamba(f la )+Mamba(f ra )
[0213] In the above formula, f Mamba represents the output of the six-way Mamba module HoM, f as Represents the multi-directional spatial feature f w Anterior flattening sequence of sagittal plane, f ps Represents the multi-directional spatial feature f w Posterior flattening sequence of sagittal plane, f sc Represents the multi-directional spatial feature f w The flattened sequence of the coronal plane, f ic Represents the multi-directional spatial feature f w The inferior flattened sequence of the coronal plane, f la Represents the multi-directional spatial feature f w The left-hand flattened sequence of the cross section, f ra Represents the multi-directional spatial feature f w Right-hand flattened sequence of cross sections.
[0214] Further, S6 comprises the following steps:
[0215] S6.1: Fuse the outputs of the adaptive gated attention module AGAM and the six-way Mamba module HoM to obtain the fused feature f merge ;
[0216] S6.2: Enhance the fused features to obtain enhanced features.
[0217] Further, S6.1 includes the following steps:
[0218] The multi-directional spatial features enhanced by the six-way Mamba attention module HoM and the adaptive gated attention module AGAM are fused to obtain the fused feature f merge , the operation can be described as:
[0219] f merge =f AGAM +f Mamba
[0220] In the above formula, f merge represents fusion features;
[0221] Further, S6.2 includes the following steps:
[0222] For the fusion feature f merge Perform layer normalization to balance the distribution of features, and then use a multi-layer perceptron to perform nonlinear transformation and feature mapping to further enhance the expressiveness of features and obtain the enhanced feature f s , the operation can be described as:
[0223] f s =MLP(LN(f merge ))+f merge
[0224] In the above formula, f s stands for enhanced features, MLP(.) stands for multi-layer perceptron, and LN(.) stands for layer normalization.
[0225] Further, refer to Figure 5 , S7 comprises the following steps:
[0226] S7.1: Calculate the channel selectivity coefficient of the enhanced feature;
[0227] S7.2: Calculate uncertainty estimates for enhanced characteristics;
[0228] S7.3: Construct channel parsing entropy module CPEM to screen and enhance features.
[0229] Further, S7.1 includes the following steps:
[0230] The input enhanced features are globally averaged and pooled, and the enhanced features of each channel are mapped to a scalar, which is then input into the multi-layer perceptron. The importance of each channel is modeled through nonlinear transformation, and the dependencies between channels are further extracted. Finally, normalization is performed to map the important coefficients of each channel to the range of [0, 1] to generate channel selection coefficients. This operation can be described as:
[0231] W = Sigmoid(MLP(GAP(f s )))
[0232] In the above formula, W represents the channel selection coefficient, and GAP(.) represents the global average pooling;
[0233] Further, S7.2 includes the following steps:
[0234] Given the current feature dimension C, calculate the channel average of the current dimension to obtain the normalized representation f of the single-scale feature sigle , the operation can be described as:
[0235]
[0236] In the above formula, represents the enhanced features of the i-th channel;
[0237] The information entropy formula is used to calculate the uncertainty estimation coefficient u of the feature. This operation can be described as:
[0238] u=-f s log(f s )
[0239] In the above formula, u is a constant, -f s log(f s ) represents information entropy, which measures the uncertainty of the feature. A larger entropy value indicates a higher uncertainty of the feature.
[0240] Further, S7.3 includes the following steps:
[0241] Adjust each channel feature, integrate the channel importance weight W and uncertainty estimation coefficient u, and obtain the filtered features. This operation can be described as:
[0242] f out =Wf s +(1-u)f s
[0243] In the above formula, f out Represents the final feature.
[0244] Further, refer to Figure 2 , S8 comprises the following steps:
[0245] S8.1: Concatenate the final features of each scale to obtain the final feature map;
[0246] S8.2: Construct a segmentation head to convert the final feature map into a segmentation map;
[0247] S8.3: Use the training set, validation set, and test set to train, validate, optimize, and test the model.
[0248] Further, S8.2 comprises the following steps:
[0249] f class =Conv(f out )
[0250]
[0251]
[0252] In the above formula, f class Represents f out The output after the convolution operation, f class [i, j, k, c] represents f class The pixel (i, j, k, c), i represents the horizontal coordinate, j represents the vertical coordinate, k represents the vertical coordinate, c represents the category of the pixel, P i,j,k,c represents the probability that the pixel (i, j, k, c) belongs to category c, argmax represents the output category label based on the maximum probability, and F[i, j, k] represents the value of the (i, j, k) pixel of the segmentation map;
[0253] Further, S8.3 includes the following steps:
[0254] Use the training set and validation set to train and optimize the model to obtain the optimal model weights, then import the optimal model weights into the test code and use the test set to test the model.
[0255] The above is a specific description of the preferred implementation of the present invention, but the invention is not limited to the embodiments. Those skilled in the art may make other equivalent modifications or substitutions without violating the spirit of the invention, and these equivalent modifications or substitutions are included in the scope defined by the application claims.
Claims
1. A liver tumor segmentation method based on multi-directional spatial features, characterized in that the method comprises the following steps: S1: Preprocess abdominal CT (Computerized Tomography) data to obtain NPY matrix data, and divide the data into training set, validation set and test set; S2: Extract the initial features of NPY matrix data; S3: Extract multi-directional spatial features from the initial features; S4: Construct an adaptive gated attention module AGAM (Adaptive Gated Attention Module) to extract multi-directional spatial features and obtain local features; S5: Construct a six-directional Mamba module HoM (Hexa-orientated Mamba) to extract multi-directional spatial features and obtain global features; S6: local features and global features are fused and enhanced to obtain enhanced features; S7: Construct a channel parsing entropy module CPEM (Channel Parsing Entropy Module) to screen the enhanced features and obtain the final features; S8: Concatenate the final features of multiple scales to obtain the segmentation map of liver tumor.
2. The method for liver tumor segmentation based on multi-directional spatial features according to claim 1, characterized in that: S1 includes the following steps: S1.1: Collect the original abdominal CT data, perform preprocessing of standardization, cropping, and data enhancement on the data, and convert it into NPY matrix data f0; S1.2: Divide the preprocessed data into training set, validation set and test set.
3. The method for liver tumor segmentation based on multi-directional spatial features according to claim 1, characterized in that: S2 includes the following steps: A deep convolution is introduced to perform a convolution operation on the NPY matrix data f0 to obtain the initial feature f. The operation can be described as: f = DepthwiseConv(f0, s, p) In the above formula, f0 represents NPY matrix data, and f=DepthwiseConv(·, s, p) represents a depthwise convolution with a stride of s and a padding size of p.
4. The method for liver tumor segmentation based on multi-directional spatial features according to claim 1, It is characterized in that S3 includes the following steps: Construct a multi-directional spatial feature extractor MDFE (Multi Directional Feature Extractor), the operation can be described as: A set of spatial separation convolutions are introduced to perform continuous convolution operations on the initial feature f along the horizontal, vertical and depth directions to obtain the multi-directional spatial feature f w , which is expressed as follows: f h =Conv 3×1×1 (f) f v =Conv 1×3×1 (f h ) f w =Conv 1×1×3 (f v ) In the above formula, f h represents the output of f after horizontal convolution, f v Represents f h The output after vertical convolution, f w represents f v The output after depthwise convolution, f h 、f v and f w Collectively called multi-directional spatial features, Conv 3×1×1 (.) represents the horizontal convolution calculation, Conv 1×3×1 (.) represents the vertical convolution calculation, Conv 1×1×3 (.) represents the depth-wise convolution calculation.
5. The method for liver tumor segmentation based on multi-directional spatial features according to claim 1, characterized in that: S4 includes the following steps: S4.1: Extract features from cross-sectional slices of multi-directional spatial features; S4.2: Calculate attention weights; S4.3: Perform local attention enhancement on multi-directional spatial features to obtain local features f AGAM ; Among them, S4.1 includes the following steps: f l =Conv(f w ) In the above formula, f l Represents cross-sectional characteristics; Among them, S4.2 includes the following steps: For multi-directional spatial features f w Perform up and down sampling and generate attention weights through convolution. The operation can be described as: f down =Downsampled(f w ) f cce =Conv(f down ) f up =Upsampled(f coe ) α=Sigmoid(f up ) In the above formula, f down Represents f w The output after downsampling operation, f coe Represents f down The output after the convolution operation, f up Represents f coe The output after upsampling operation, α represents the attention weight, Downsampled(.) represents downsampling, which means reducing the spatial resolution of the input, Upsampled(.) represents upsampling, which means increasing the spatial resolution of the input, and Sigmoid(.) represents normalization; Among them, S4.3 includes the following steps: The attention weight α is used to adjust the multi-directional spatial features and fuse them with the cross-sectional features fl. This operation can be described as: f AGAM =f l +αf w In the above formula, αf w represents the use of attention weights to adjust multi-directional spatial features, f AGAM Represents the output of the adaptive gated attention module AGAM.
6. The method for liver tumor segmentation based on multi-directional spatial features according to claim 1, characterized in that: S5 includes the following steps: S5.1: Construct Mamba long sequence modeling operation; S5.2: Flatten the coronal, sagittal and transverse sequences of multi-directional spatial features into Mamba long sequence modeling; Among them, S5.1 includes the following steps: Let the dimension of the feature be t, and define f t ={x1, x2, ..., x t }, f′ t ={y′1,y′2,...,y′ t }, f′ t =Mamba(f t ); Among them, x t 2D flattened sequence representing features, f t Representative feature, f′ t Representative feature f t The output after Mamba long sequence modeling, y′ t represents f′ t 2D flattened sequence, Mamba(.) represents the Mamba long sequence modeling operation; The Mamba long sequence modeling operation Mamba(.) is defined as follows: x t After linear transformation and convolution operation, the expression is: x t =Conv(Linear(x t )) In the above formula, Linear(.) represents linear transformation; Introducing parameter B t and C t , whose expression is: B t =Linear B (x t ) C t =Linear C (x t ) In the above formula, B t Represents according to x t The dynamically generated input transformation matrix, C t Represents according to x t Dynamically generated input transformation matrix, Linear B (.) stands for B t Linear transformation, Linear C (.) stands for C t Linear transformation of ; Construct the state space model, which is expressed as: h t =Ah t-1 +B t x t y t =C t h t In the above formula, h t represents the hidden state of the state space model, y t represents the predicted output of the state space model, and A represents the state transfer matrix; Based on the state space model, the dynamic time step Δ is introduced t , whose expression is: Δ t =Linear Δ (x t ) In the above formula, Linear Δ (.) represents a special case for Δ t Linear transformation of ; Construct a time-varying state space model, which is expressed as: A′=exp(Δ t A) B′=(Δ t A) -1 (exp(Δ t A)-I)B t In the above formula, A′, B′ represent the time step Δ t After discretization, I represents the diagonal matrix, and exp(.) represents the natural exponential function; h′ t =A′h′ t-1 +B′x t y′ t =C t h′ t In the above formula, h′ t represents the hidden state of the time-varying state space model, y′ t represents the predicted output of the time-varying state-space model; The multi-directional spatial feature f w After Mamba long sequence modeling operation, we can get f′ w , whose expression is: f′ w =Mamba(f w ) In the above formula, f′ w Represents the multi-directional spatial feature f w Output after Mamba long sequence modeling; Among them, S5.2 includes the following steps: The multi-directional spatial features are flattened into 6 feature sequences, and Mamba long sequence modeling is performed on these 6 sequences respectively. Then, the modeled sequences are fused to achieve global attention enhancement of the multi-directional spatial features. This operation can be described as: f Mamba =Mamba(f as )+Mamba(f ps )+Mamba(f sc )+Mamba(f ic )+Mamba(f la )+Mamba(f ra ) In the above formula, f Mamba represents the output of the six-way Mamba module HoM, f as Represents the multi-directional spatial feature f w Anterior flattening sequence of sagittal plane, f ps Represents the multi-directional spatial feature f w Posterior flattening sequence of sagittal plane, f sc Represents the multi-directional spatial feature f w The flattened sequence of the coronal plane, f ic Represents the multi-directional spatial feature f w The inferior flattened sequence of the coronal plane, f la Represents the multi-directional spatial feature f w The left-hand flattened sequence of the cross section, f ra Represents the multi-directional spatial feature f w Right-hand flattened sequence of cross sections.
7. The method for liver tumor segmentation based on multi-directional spatial features according to claim 1, characterized in that: S6 includes the following steps: S6.1: Fuse the outputs of the adaptive gated attention module AGAM and the six-way Mamba module HoM to obtain the fused feature f merge ; S6.2: Enhance the fused features to obtain enhanced features; Wherein, S6.1 comprises the following steps: The multi-directional spatial features enhanced by the six-way Mamba attention module HoM and the adaptive gated attention module AGAM are fused to obtain the fused feature f merge , the operation can be described as: f merge =f AGAM +f Mamba In the above formula, f merge represents fusion features; Among them, S6.2 includes the following steps: For the fusion feature f merge Perform layer normalization to balance the distribution of features, and then use a multi-layer perceptron to perform nonlinear transformation and feature mapping to further enhance the expressiveness of features and obtain the enhanced feature f s , the operation can be described as: f s =MLP(LN(f merge ))+f merge In the above formula, f s stands for enhanced features, MLP(.) stands for multi-layer perceptron, and LN(.) stands for layer normalization.
8. The method for liver tumor segmentation based on multi-directional spatial features according to claim 1, characterized in that: S7 includes the following steps: S7.1: Calculate the channel selectivity coefficient of the enhanced feature; S7.2: Calculate uncertainty estimates for enhanced characteristics; S7.3: Construct the channel parsing entropy module CPEM (Channel Parsing Entropy Module) to screen and enhance features; Among them, S7.1 includes the following steps: The input enhanced features are globally averaged and pooled, and the enhanced features of each channel are mapped to a scalar, which is then input into the multi-layer perceptron. The importance of each channel is modeled through nonlinear transformation, and the dependencies between channels are further extracted. Finally, normalization is performed to map the important coefficients of each channel to the range of [0, 1] to generate channel selection coefficients. This operation can be described as: W=Sigmoid(MLP(GAP(f s ))) In the above formula, W represents the channel selection coefficient, and GAP(.) represents the global average pooling; Among them, S7.2 includes the following steps: Given the current feature dimension C, calculate the channel average of the current dimension to obtain the normalized representation f of the single-scale feature sigle , the operation can be described as: In the above formula, represents the enhanced features of the i-th channel; The information entropy formula is used to calculate the uncertainty estimation coefficient u of the feature. This operation can be described as: u=-f s log(f s ) In the above formula, u is a constant, -f s log(f s ) represents information entropy, which measures the uncertainty of the feature. A larger entropy value indicates a higher uncertainty of the feature. Among them, S7.3 includes the following steps: Adjust each channel feature, integrate the channel importance weight W and uncertainty estimation coefficient u, and obtain the filtered features. This operation can be described as: f out =Wf s +(1-u)f s In the above formula, f out Represents the final feature.
9. The method for liver tumor segmentation based on multi-directional spatial features according to claim 1, characterized in that: S8 includes the following steps: S8.1: Concatenate the final features of each scale to obtain the final feature map; S8.2: Construct a segmentation head to convert the final feature map into a segmentation map; S8.3: Use the training set, validation set, and test set to train, validate, optimize, and test the model; Among them, S8.2 includes the following steps: f class =Conv(f out ) In the above formula, f class Represents f out The output after the convolution operation, f class [i, j, k, c] represents f class The pixel (i, j, k, c), i represents the horizontal coordinate, j represents the vertical coordinate, k represents the vertical coordinate, c represents the category of the pixel, P i,j,k,c represents the probability that the pixel (i, j, k, c) belongs to category c, argmax represents the output category label based on the maximum probability, and F[i, j, k] represents the value of the (i, j, k) pixel of the segmentation map.
Citation Information
Patent Citations
Brain tumor image segmentation method based on multi-scale convolution and Mama structure
CN118447244A
Three-dimensional oral hard palate image segmentation method based on multidirectional state space model
CN118941585A
Medical image segmentation method based on global and local feature reconstruction network
WO2023151141A1