A medical image segmentation method based on adaptive anisotropic convolution

By using an adaptive anisotropic convolutional layer and a cross-scale feature fusion three-dimensional medical image segmentation network, the problem of global information capture and detail preservation in kidney tumor segmentation using deep learning methods has been solved, achieving high-precision segmentation of small-scale kidney tumors and improving the accuracy and efficiency of surgery.

CN120726076BActive Publication Date: 2025-11-07NANCHANG CAMPUS OF EAST CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511244241.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-11-07
Estimated Expiration
2045-09-02

AI Technical Summary

Technical Problem

Existing deep learning methods struggle to effectively capture global information in kidney tumor segmentation tasks, especially for small tumors with complex shapes and detailed boundaries, where segmentation accuracy is insufficient. Furthermore, their performance is limited under multi-scale feature adaptation and noise interference.

Method used

An adaptive anisotropic convolutional layer is employed, and a three-dimensional medical image segmentation network is designed by combining cross-scale feature fusion and multi-stage deep supervision through parallel multimodal convolution and adaptive attention mechanism to optimize feature extraction and segmentation accuracy.

Benefits of technology

It significantly improves the segmentation accuracy of small-scale targets, especially the segmentation effect of complex organs such as kidney tumors, simplifies the surgical procedure, and improves the applicability and accuracy of the segmentation model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726076B_ABST
    Figure CN120726076B_ABST
Patent Text Reader

Abstract

The application provides a medical image segmentation method based on adaptive anisotropic convolution, comprising: acquiring a three-dimensional medical CT data set containing multiple abdominal organs and kidney tumors and labels, and preprocessing the data set; dividing the data set into a training set and a test set for model training and evaluation; designing a three-dimensional medical image segmentation network model based on an adaptive anisotropic convolution layer, inputting the preprocessed training set into the three-dimensional medical image segmentation network model, training the three-dimensional medical image segmentation network model through parallel multi-modal convolution, adaptive attention weight generation, weighted feature dynamic fusion and multi-stage deep supervision, and optimizing model parameters; and applying the optimized three-dimensional medical image segmentation network model to the test set to generate a three-dimensional segmentation result with clear boundaries and complete details, thereby providing support for clinical diagnosis and treatment planning.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of information technology, and in particular to a medical image segmentation method based on adaptive anisotropic convolution. BACKGROUND

[0002] In recent years, with the continuous development of modern medical technology, the diagnosis technology of various kidney diseases has made significant progress, but kidney cancer is still one of the ten most common malignant tumors in the world, which seriously threatens human health. In particular, the incidence of small kidney cancer is rising, and how to achieve precise treatment has become the focus of clinical attention. Partial nephrectomy (PN) as a minimally invasive surgical procedure can remove kidney tumors while preserving as much normal kidney function as possible, and has become the preferred method for the treatment of small kidney cancer. In particular, laparoscopic partial nephrectomy (LPN) and robotic-assisted partial nephrectomy (RAPN) have been widely used and played a crucial role in clinical practice due to their advantages of precise resection and minimal damage to normal tissue. LPN and RAPN mainly remove the diseased part of the kidney while preserving normal kidney tissue with good function, greatly optimizing the treatment effect.

[0003] Based on the deep learning-based kidney tumor segmentation technology, it can effectively assist the identification of tumor location in surgery and provide precise reference for clinical treatment. However, although deep learning methods have shown strong capabilities in the field of medical image segmentation, there are still many challenges for this complex target, such as the diversity of tumor size and shape, the difficulty of distinguishing from normal tissue, and the noise interference of image quality (such as uneven brightness or low contrast) on model input.

[0004] Currently, deep learning technology has gradually become the mainstream method of medical image segmentation and has improved the accuracy and efficiency of medical image segmentation. Deep learning-based medical image segmentation methods can be mainly divided into the following four categories: convolutional neural network (CNN) based methods, Transformer-based methods, hybrid architecture based on CNN and Transformer, and state space model (SSM) based methods, specifically:

[0005] 1) CNN-based methods usually use powerful CNN backbone networks, such as Res2Net (Residual 2-stream Network) and HardNet68 (Harmonious Dense Network 68), combined with enhancement modules and operations to optimize feature extraction. However, the limitations of convolutional operations make it difficult to capture global information, especially for complex shape kidney tumor segmentation tasks. In addition, the inductive bias of CNN reduces its adaptability in processing laparoscopic images with variable structures;

[0006] 2) Transformers excel at handling long-range dependencies through their core self-attention mechanism. For example, Vision Transformer (ViT) was the first to apply the Transformer architecture to computer vision tasks. However, existing Transformer-based methods may still lose small tumors and partial boundary details during the decoding stage, and have limited ability to adapt to multi-scale features;

[0007] 3) To better extract global semantic features and local detail information, hybrid network segmentation models combine the advantages of CNN and Transformer. However, serial hybrid methods do not fully utilize the advantages of multi-scale feature fusion by sequentially stacking CNN and Transformer blocks. Parallel structures, while overcoming the limitations of serial models, have large parameter quantities and limited fusion effects, while introducing background noise interferes with segmentation performance;

[0008] 4) Based on state space models (SSM), such as the Mamba (Memory-augmented Model with Bidirectional Attention) model, structured state space blocks are used to capture long-range dependencies while maintaining linear computational complexity. Some models, such as VMambaUNet (Vision Mamba U-shaped Network), introduce visual state space modules to further optimize segmentation performance. However, these methods are still limited by dataset characteristics and have weak generalization ability.

[0009] In addition, attention mechanisms are crucial in medical image segmentation. Their main approaches include activation weight-based attention mechanisms and non-local interaction attention mechanisms. Single attention designs are simple and easy to implement but have insufficient performance; serial attention has high computational complexity; parallel attention has higher efficiency but may cause feature redundancy. How to effectively guide the model to focus on key feature regions to achieve fine-grained segmentation is still a technical challenge that needs to be addressed in the field of medical image segmentation.

[0010] In summary, for the complex kidney tumor image segmentation task, how to design a multi-modal, cross-scale deep learning method to accurately locate the tumor position while achieving more fine-grained prediction, reducing the complexity of the surgeon's operation and improving the treatment effect is the core problem that needs to be solved by those skilled in the art. SUMMARY

[0011] The application provides a medical image segmentation method based on adaptive anisotropic convolution. Specifically, the application provides a new convolution operation mechanism and network architecture to improve the feature extraction capability of anisotropic CT images and enhance the segmentation accuracy of small-scale targets.

[0012] The first aspect of the application provides a medical image segmentation method based on adaptive anisotropic convolution, comprising:

[0013] S101: Obtain a three-dimensional medical CT data set containing multiple abdominal organs and kidney tumors and labels, and pre-process the data set;

[0014] S102: Divide the data set into a training set and a test set for model training and evaluation;

[0015] S103: Design a three-dimensional medical image segmentation network model based on an adaptive anisotropic convolution (Adaptive Asymmetric Convolution, AAs-conv) layer, input the pre-processed training set into the three-dimensional medical image segmentation network model, and train the three-dimensional medical image segmentation network model through parallel multi-modal convolution, adaptive attention weight generation, weighted feature dynamic fusion, and multi-stage deep supervision to optimize the model parameters;

[0016] S104: Apply the optimized three-dimensional medical image segmentation network model to the test set to generate a three-dimensional segmentation result with clear boundaries and complete details.

[0017] Further, the data set pre-processing includes:

[0018] CT value truncation: The CT value of the abdominal CT image containing the fine annotation of the kidney and the kidney tumor is truncated in the range of [-75, 293] Hu (Hounsfield unit), wherein -75 and 293 are the 0.5% quantile and the 99.5% quantile of the foreground region CT value, respectively, to reduce the influence of irrelevant background regions;

[0019] Resampling: According to the median of the data set voxel spacing, set the target spacing, resample the CT image to [3.22mm, 1.62mm, 1.62mm], and unify the voxel spacing of each case image to standardize the input voxel block, facilitating batch training of the model;

[0020] Standardization: Z-score standardization method is adopted, that is, any sample value x in the data set is subtracted from the mean of the sample and divided by the standard deviation of the sample, so that the resampled image data is unified to a standard distribution with a mean of 0 and a standard deviation of 1. The CT image of the experimental data KiTS19 in this chapter has a mean of 105.0 and a variance of 73.9.

[0021] Further, the three-dimensional medical image segmentation network model of adaptive anisotropic convolution comprises:

[0022] Adaptive anisotropic convolution (AAs-conv) layer: simultaneously performing standard three-dimensional convolution 3x3x3 and spatial separable convolution 3x1x1+1x3x3 dual-path operation, dynamically fusing two feature maps through dynamic weight mechanism, dynamically generating by channel attention mechanism, and adaptively balancing spatial information interaction and anisotropic feature decoupling;

[0023] Symmetric encoding and decoding structure: 6 layers of encoder, channel number from 24 to 320, 5 layers of decoder, channel number from 320 to 24, and segmentation result of each layer of decoder is output for deep supervision; wherein, the encoding path is composed of 6 encoder modules, each module contains two AAs-conv layers, and down-sampling is performed through convolution operations with different steps to generate feature maps layer by layer, and the channel number of the feature maps increases layer by layer and the resolution decreases layer by layer; the decoding path is composed of 5 decoder modules, each module contains an AAs-conv layer, and the resolution of the feature map is restored through transposed convolution for up-sampling, and the feature map is spliced with the feature map of the corresponding layer of the encoder to fuse the shallow details and deep semantic information;

[0024] Cross-scale feature fusion module: in the encoding path, the original resolution feature map output by the first layer of the encoder is added to the deep low-resolution feature map after average pooling and convolution scaling, to enhance the detail retention ability of small-scale targets;

[0025] Further, the deep supervision mechanism outputs a segmentation result at each layer (except the bottom layer) of the decoder through a 1x1x1 convolution bypass, and calculates the loss to supervise different depths of the network and relieve the gradient vanishing problem.

[0026] Further, the AAs-conv layer comprises:

[0027] Dual-path feature extraction module: receiving input feature maps to perform dual-path parallel convolution operation, extracting two types of spatial features through the spatial separable convolution path 3x1x1-1x3x3 of the first branch and the standard three-dimensional convolution path 3x3x3 of the second branch;

[0028] The attention weight generation module integrates the output features of the spatial separation convolution and the three-dimensional spatial convolution into a global feature map, compresses the spatial information of each channel of the global feature map by using average pooling, obtains the attention vectors of the branch information after convolution compression of the channel size and convolution recovery, splices the attention vectors of the branch information in the channel dimension, and then performs softmax mapping on the branch information in the channel dimension to obtain adaptive weights of the two feature maps;

[0029] The weighted fusion module weights and fuses the feature maps of the two paths according to the weights of the two branches, and outputs the fused feature maps.

[0030] Further, the cross-scale feature fusion operation is:

[0031] In the encoding path, the original resolution feature map output by the first layer encoder is down-sampled to the same resolution as the current layer feature map by average pooling, and the resolutions of different layers are unified by converting the resolution.

[0032] The down-sampled feature map is added and fused with the deep low-resolution feature map after being scaled and adjusted in channel number by 1x1x1 convolution;

[0033] Specifically, it is and wherein represents the output of the first layer encoder, represents the output before fusion of the corresponding layer encoder, and is greater than 1, represents the output after fusion of the corresponding layer encoder, and is greater than 1, represents the cross-scale connection mapping, represents the LeakyReLU nonlinear transformation operation, represents the average pooling operation, represents the convolution operation with a kernel of 1x1x1, represents the InstanceNorm normalization operation.

[0034] Further, the training process includes:

[0035] Input data cropping: the preprocessed three-dimensional abdominal CT image is cropped into a fixed size voxel block (64X128X128) as the network input, and different sub-regions are covered by a random center point cropping strategy;

[0036] Online data augmentation: During training, an online data augmentation strategy is adopted to expand the training samples, including random rotation within the range of [-30°, 30°], random size scaling of 0.7 times to 1.4 times, random flipping, random brightness transformation, random contrast transformation, random gamma transformation, and various transformations combined randomly.

[0037] Multi-stage deep supervision: At the end of each layer of the 5-layer decoder, a 1X1X1 convolutional layer and a softmax activation function are used to output a segmentation prediction map of the corresponding resolution. The composite loss function is calculated for each layer prediction map.

[0038] Further, the loss function is:

[0039] The formula is Wherein represents the cross-entropy loss function, represents the number of sample pixels, and respectively represent the values of the prediction result and the true label at the i pixel point, represents the natural logarithm of the probability of predicting a positive class, which is used to measure the confidence of predicting a positive class, represents the natural logarithm of the probability of predicting a negative class, which is used to measure the confidence of predicting a negative class.

[0040] The loss function is not only used to measure the difference between the final prediction result and the true label, but also serves as the core driving force of the multi-stage deep supervision strategy. By calculating the loss at each decoding stage and performing weighted summation, the semantic constraints of shallow, middle and deep features are obtained, thereby significantly improving the segmentation accuracy, accelerating the network convergence speed, and improving the detection ability of small voxel proportion targets.

[0041] The second aspect of the present application provides a kidney and kidney tumor diagnosis and treatment auxiliary equipment, comprising a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to execute the steps of the medical image segmentation method based on adaptive anisotropic convolution as described above.

[0042] Compared with the prior art, the present application has the following beneficial effects:

[0043] 1) The present application introduces an adaptive anisotropic convolution layer, which extracts different directional feature information in three orthogonal directions in parallel, and dynamically allocates weights to each branch using an adaptive attention mechanism, achieving accurate fusion of directional features and significantly improving the ability to capture different directional structural information, especially in the segmentation of complex-shaped small-scale targets such as kidney tumors.

[0044] 2) The present application significantly improves the segmentation accuracy of small-scale organs with fuzzy boundaries such as the duodenum, gallbladder, prostate, and uterus. Experimental results show that the mDice (average Dice coefficient, an evaluation index for measuring the similarity between the segmentation results of the algorithm and the true label) of the three-dimensional medical image segmentation network model with adaptive anisotropic convolution layer is 90.07%, and the mNSD (Modified Normalized Surface Dice) is 79.04%, which is better than the advanced model such as nnUNet (No New-Net, an adaptive deep learning framework), and the Dice of the duodenum and gallbladder is increased by 1.26% and 2.50% respectively. In the non-significant anisotropic scene (such as KiTS19 kidney tumor segmentation), the AAs-conv adaptively fuses the advantages of spatial separation convolution and standard three-dimensional convolution, and combines with cross-scale fusion to retain details, which significantly improves the segmentation accuracy of small-scale and complex texture kidney tumors (in the experiment, the Dice of the three-dimensional medical image segmentation network model with adaptive anisotropic convolution layer for kidney tumor is 85.10%, and the Dice of the kidney is 97.04%, which is increased by 3.30% and 1.19% respectively compared with the classic 3D U-Net);

[0045] 3) The method of the present application is a single-stage segmentation model, and the processing flow is simple and efficient, which is suitable for clinical practice application;

[0046] 4) The present application has wide applicability, and the method can be effectively applied to various medical CT image segmentation tasks, especially in the scene involving small-scale, fuzzy boundary targets (organs or lesions) and anisotropic images, such as abdominal multi-organ segmentation, kidney and kidney tumor segmentation, etc., which provides reliable technical support for precise medical diagnosis and surgical planning. BRIEF DESCRIPTION OF DRAWINGS

[0047] Figure 1 A flow chart of a medical image segmentation method based on adaptive anisotropic convolution provided by an embodiment of the present application.

[0048] Figure 2 A network structure diagram of a three-dimensional medical image segmentation network model with adaptive anisotropic convolution provided by an embodiment of the present application.

[0049] Figure 3 A network structure diagram of an adaptive anisotropic convolution (AAs-conv) layer provided by an embodiment of the present application. DETAILED DESCRIPTION

[0050] The embodiment of the present application provides a medical image segmentation method based on adaptive anisotropic convolution, which is used to solve the technical problem that the segmentation accuracy of the existing three-dimensional convolutional neural network is limited when processing medical CT images with anisotropy (inconsistent scanning interlayer and intralayer voxel spacing), especially for small-scale, boundary fuzzy organs (such as the duodenum and gallbladder) or complex texture small-scale lesions (such as kidney tumors).

[0051] In order for those skilled in the art to better understand the technical solutions in the specification, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments. Obviously, the described embodiments are only a part of the embodiments of the specification, not all the embodiments. Based on the embodiments in the specification, all other embodiments obtained by those skilled in the art without creative labor should belong to the protection scope of the specification.

[0052] Please refer to Figure 1 , Figure 1 The step flow chart of the medical image segmentation method based on adaptive anisotropic convolution provided by the embodiment of the present application is shown in the following figure.

[0053] The present application provides a medical image segmentation method based on adaptive anisotropic convolution, comprising:

[0054] Step S101: Obtain a three-dimensional medical CT data set containing multiple abdominal organs and kidney tumors and image and label, and pre-process the data set;

[0055] Step S102: Divide the data set into a training set and a test set for model training and evaluation;

[0056] Step S103: Design a three-dimensional medical image segmentation network model based on an adaptive anisotropic convolution (Adaptive Asymmetric Convolution, AAs-conv) layer, input the pre-processed training set into the three-dimensional medical image segmentation network model, train the model by parallel multi-modal convolution, adaptive attention weight generation, weighted feature dynamic fusion and multi-stage deep supervision, and optimize the model parameters;

[0057] Step S104: Apply the optimized three-dimensional medical image segmentation network model to the test set to generate a three-dimensional segmentation result with clear boundaries and complete details.

[0058] In one specific embodiment of the present embodiment, the source of the CT image data set is:

[0059] The present embodiment uses the KiTS19 (Kidney Tumor Segmentation 2019) public dataset for training and testing; the dataset contains 300 abdominal CT images and their corresponding kidney and kidney tumor fine annotations, which can fully cover the diversity characteristics of the kidney and tumor. Among them, 210 cases are randomly divided into training set and validation set according to the ratio of 8:2, and the remaining 90 cases are used as the test set.

[0060] In one specific embodiment of the present embodiment, the preprocessing process of the CT image data includes:

[0061] CT value truncation (Windowing): Obtain the pixel gray value (i.e. CT value) of the CT image, and truncate the CT value in the preset range [-75, 293] Hu (Hounsfield unit), highlight the soft tissues such as kidney and tumor, and the range is calculated by the 0.5% quantile and 99.5% quantile of the foreground region CT value in the training data, which effectively suppresses the interference of irrelevant background noise;

[0062] Resampling of voxels: Set the target interval according to the median of the data set voxel interval, and the median resolution of the resampled image is 127x247x247. Resample all CT images to a uniform target voxel interval, such as [3.22mm, 1.62mm, 1.62mm], to unify the physical scale of different cases, standardize the image input size, and facilitate batch training of the model;

[0063] Data normalization: Z-score normalization method is used, that is, any sample value in the data set is subtracted from the sample mean and divided by the sample standard deviation. The resampled image data is standardized to a standard distribution with a mean of 0 and a standard deviation of 1, which accelerates model convergence and improves training stability. The CT image mean of the experimental data KiTS19 is 105.0, and the variance is 73.9.

[0064] In one specific embodiment of the present embodiment, the training process of the model includes:

[0065] Input data cropping: The preprocessed three-dimensional abdominal CT image is cropped into a fixed size voxel block (64X128X128) as network input, and different sub-regions are covered by random center point cropping strategy;

[0066] Online data augmentation: online data augmentation strategy is adopted during training to expand training samples, including random rotation within [-30°, 30°] range, random size scaling of 0.7 times to 1.4 times, random flipping, random brightness transformation, random contrast transformation, random gamma transformation, and various transformations of random combination of the above;

[0067] Multi-stage deep supervision: at the end of each layer of the 5-layer decoder, a 1x1x1 convolutional layer and a softmax activation function are used to output a segmentation prediction map of corresponding resolution, and a composite loss function is calculated for each layer prediction map, and the loss function is wherein represents the cross-entropy loss function, represents the number of sample pixels, , and respectively represent the values of the prediction result and the true label at the i th pixel point, represents the natural logarithm of the probability of predicting positive class, which is used to measure the confidence of predicting positive class, represents the natural logarithm of the probability of predicting negative class, which is used to measure the confidence of predicting negative class. The loss function is not only used to measure the difference between the final prediction result and the true label, but also serves as the core driving force of the multi-stage deep supervision strategy. By calculating the loss at each decoding stage and performing weighted summation, the semantic constraints of shallow, middle and deep features are obtained, thereby significantly improving the segmentation accuracy, accelerating the network convergence speed, and improving the detection ability of small voxel proportion targets.

[0068] In one specific embodiment of the present embodiment, the three-dimensional medical image segmentation network with adaptive anisotropic convolution (see Figure 2 ) is composed of a model, which includes:

[0069] Adaptive anisotropic convolution (AAs-conv) layer: simultaneously performing standard three-dimensional convolution (3x3x3) and spatially separated convolution (3x1x1+1x3x3) dual-path operation, fusing two feature maps through dynamic weight mechanism, dynamically generated by channel attention mechanism, and adaptively balancing spatial information interaction and anisotropic feature decoupling;

[0070] Symmetric encoding and decoding structure: 6 layers of encoder (channel number 24→320) and 5 layers of decoder (channel number 320→24), and each layer of decoder outputs a segmentation result for deep supervision;

[0071] In the encoding path, six encoder modules are included, each of which contains two AAs-conv layers, and down-sampling is performed through convolution operations with different steps (cross-step convolution) to generate feature maps with more channels and lower resolution layer by layer; in the decoding path, five decoder modules are included, each of which contains one AAs-conv layer, and up-sampling is performed through transposed convolution to restore the resolution of the feature maps, and the feature maps after the transposed convolution up-sampling are spliced with the feature maps output by the corresponding layer of the encoder through the way of jump connection to fuse the shallow details and deep semantic information; the deep supervision mechanism bypasses a 1x1x1 convolution to output a segmentation result at each layer of the decoder (except the bottom layer), and calculates the loss, thereby supervising different depths of the network and relieving the gradient vanishing problem.

[0072] The cross-scale feature fusion module: in the encoding path, the original resolution feature map output by the first layer of the encoder is added and fused with the deep low-resolution feature map after being scaled through average pooling and convolution, thereby enhancing the detail retention ability of small-scale targets.

[0073] In one specific embodiment of the present embodiment, the AAs-conv layer is composed of:

[0074] The double-path feature extraction module: receives the input feature map, and extracts two types of spatial features through a spatial separable convolution path (3x1x1-1x3x3) and a standard three-dimensional convolution path (3x3x3), respectively.

[0075] The attention weight generation module: performs global pooling, channel compression and Softmax processing on the double-path output features to generate channel-level weight coefficients for each path.

[0076] The weighted fusion module: the feature maps of the two paths are weighted and summed according to the attention weights, and the fused feature map is output.

[0077] In one specific embodiment of the present embodiment, the adaptive anisotropic convolution (AAs-conv) layer basic feature extraction unit includes the following steps in the feature extraction process of the encoder and the decoder:

[0078] The adaptive anisotropic convolution (AAs-conv) block performs double-path parallel operation, the first path performs the above spatial separable convolution (3x1x1 convolution+1x3x3 convolution), the second path performs standard 3x3x3 three-dimensional convolution, and the two paths are dynamically fused through the channel attention mechanism, the initial features after fusion are compressed through global average pooling, the adaptive weight vectors of the two branches are generated through the fully connected layer containing the bottleneck structure (such as dimension reduction and then dimension increase), and the spatial separable convolution contribution weight of the two feature maps is obtained by applying the Softmax function and the standard three-dimensional convolution contribution weight final weights Fusing two outputs. In addition, cross-scale feature fusion is introduced in the encoding path: the original resolution feature map output by the first layer encoder is down-sampled to the same resolution as the current layer feature map through average pooling, and then adjusted in channel number through 1x1x1 convolution, and finally added and fused with the feature map output by the current layer encoder to alleviate the loss of details in the deep down-sampling process, especially for small scale targets (such as kidney tumors).

[0079] In one specific embodiment of the present embodiment, the adaptive anisotropic convolution (AAs-conv) block adopts a spatial separation convolution strategy, including:

[0080] The features between scanning layers (usually the vertical axis z direction) are processed separately using a convolution kernel size of 3x1x1, and the features within the scanning layer (usually the transverse x-y direction) are processed separately using a convolution kernel size of 1x3x3, and finally the features are integrated and the channel is adjusted through 1x1x1 convolution. At the same time, after each upsampling operation in the decoding path, a three-way Mamba non-local operation module (Tri-directional Mamba Non-Local Operation Module, TMNLOM) is introduced, which first performs global information aggregation (for example, using global maximum pooling to compress into a sequence) along the depth (S), height (H), and width (W) directions of the feature map, then uses an efficient Mamba state space model (SSM) to model the long-distance dependency relationship of the aggregated sequence in each direction, and finally applies the global response factor obtained by modeling in the three directions to the original feature map, realizing efficient non-local global information interaction.

[0081] In one specific embodiment of the present embodiment, the adaptive anisotropic convolution (AAs-conv) layer solves the anisotropic step of the CT image in the vertical axis (Z axis) and the transverse plane (XY plane) (see Figure 3 ), including:

[0082] 1. Parallel multi-modal convolution:

[0083] The input feature map is subjected to two-path convolution operation. One branch is a 3x1x1-1x3x3 spatial separation convolution, as shown in equation (1), and the other branch is a 3x3x3 classical three-dimensional spatial convolution, as shown in equation (2). The spatial separation convolution path (3x1x1-1x3x3) extracts anisotropic direction features by decoupling the convolution operation of the Z-axis (inter-layer direction) and the XY plane (intra-layer direction), and focuses on solving the geometric distortion problem caused by low inter-layer resolution of CT images. The output feature map focuses on position accuracy (boundary positioning of small tumors). The standard three-dimensional convolution path (3x3x3) uniformly processes three-dimensional space and extracts global spatial context features (organ overall shape and tumor-tissue topological relationship). The output feature map focuses on semantic integrity.

[0084] (1)

[0085] (2)

[0086] In the formula: represents the input of AAs-conv, and respectively represent the outputs of the corresponding spatial separation convolution and three-dimensional spatial convolution (including abstract representations such as the position, edge (affecting clarity perception), texture, shape (graphical features) of objects in the image), and respectively represent LeakyReLU (Leaky Rectified Linear Unit, an improved ReLU activation function) nonlinear transformation operation and InstanceNorm (Instance Normalization) normalization operation, represents a convolution operation with a 3x1x1 kernel, represents a convolution operation with a 1x3x3 kernel, represents a convolution operation with a 3x3x3 kernel.

[0087] 2. Adaptive attention weight generation, as shown in equation (3):

[0088] (3)

[0089] In the formula: represents the integrated feature map, and respectively represent the outputs of the corresponding spatial separation convolution and three-dimensional spatial convolution;

[0090] Global average pooling is used to compress the spatial information of each channel in the integrated feature map, as shown in equation (4):

[0091] (4)

[0092] , , , respectively represent the respective direction dimensions of the corresponding input feature map, represents an average pooling function; after a 1x1x1 convolution compression channel and then through a 1x1x1 convolution to restore the channel scale, the attention vector of each branch information is obtained and , see formula (5) and (6):

[0093] (5)

[0094] (6)

[0095] , and represent the attention vector of each branch information obtained; represents an activation function for nonlinear transformation; (Instance Normalization) represents a normalization method for small batch training data, which stabilizes and speeds up the training of neural networks; represents a convolution method, which performs point convolution on volumetric data through 1x1x1 convolution operation, which is used to change the dimension of the channel and the aggregation of the feature; represents a globally averaged feature map, which is compressed in channel after a 1x1x1 convolution and then restored in channel scale through a 1x1x1 convolution;

[0096] The attention vectors of each branch information are spliced in the channel dimension, and then the softmax mapping of each branch information in the channel dimension is obtained to obtain the adaptive weight of the two feature maps and see formula (7) and (8), wherein represents the output probability of class a, represents the output probability of class b, that is, different attention is given to different branch information at each channel scale:

[0097] (7)

[0098] (8)

[0099] , and respectively represent the weight or normalization value of the different branches of the input feature; and respectively, are transformed by exponential function and respectively, to make larger values occupy a larger proportion in the normalization process.

[0100] 3. Weighted feature dynamic fusion:

[0101] According to the weights of two branches and , the output feature maps of two convolutional branches are fused to realize dynamic selection of the results of spatial separable convolution and three-dimensional spatial convolution operation, as shown in equations (9) and (10):

[0102] (9)

[0103] (10)

[0104] In the formula, represents the output feature, which is obtained by weighted summation of the input features; and respectively represent a certain weight or normalized value of different branches of the input feature; and respectively represent the input features and ( represent the weights of the features extracted by the spatial separable convolution path, represent the weights of the features extracted by the standard 3D convolution path) respectively corresponding to the weights and adjusted values, which determine the influence degree of each input feature in the final output; It is shown that the two weights form a probability distribution, the sum is 1, which indicates that the contribution of the input feature is normalized to 1 in an uncertain case, ensuring the stability and rationality of the output feature.

[0105] The embodiment of the application provides a kidney and kidney tumor segmentation auxiliary device, which comprises a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to execute the kidney and kidney tumor segmentation method based on the adaptive anisotropic convolution of the medical image segmentation network.

[0106] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described device and product can refer to the corresponding process in the foregoing method embodiment, which will not be repeated here.

[0107] In several embodiments provided in the application, it should be understood that the disclosed method, system, device and program product can be implemented in other ways. ​

[0108] In addition, each functional unit in each of the embodiments of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0109] The above description and the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit the same; although the foregoing embodiments of the present application are described in detail, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A medical image segmentation method based on adaptive anisotropic convolution, characterized in that, The application relates to a three-dimensional medical image segmentation method based on adaptive anisotropic convolution layers. The application comprises the following steps: An image and label containing a plurality of abdominal organs and kidney tumors are acquired, and a three-dimensional medical CT data set is preprocessed; The data set is divided into a training set and a test set for model training and evaluation; A three-dimensional medical image segmentation network model based on adaptive anisotropic convolution layers is designed, the preprocessed training set is input into the three-dimensional medical image segmentation network model, the three-dimensional medical image segmentation network model is trained through parallel multi-modal convolution, adaptive attention weight generation, weighted feature dynamic fusion and multi-stage deep supervision, and model parameters are optimized; The three-dimensional medical image segmentation network model based on adaptive anisotropic convolution layers comprises the following steps: An adaptive anisotropic convolution layer: simultaneously performing standard three-dimensional convolution 3x3x3 and spatially separated convolution 3x1x1+1x3x3 dual-path operation, dynamically fusing two feature maps through a dynamic weight mechanism, and dynamically generating, through a channel attention mechanism, adaptive balance of spatial information interaction and anisotropic feature decoupling; A symmetric encoding and decoding structure: an encoder with 6 layers, the channel number being increased from 24 to 320, and a decoder with 5 layers, the channel number being decreased from 320 to 24, and a segmentation result being output by each layer of the decoder for deep supervision; The cross-scale feature fusion module: in the encoding path, the original resolution feature map output by the first layer encoder is fused with the deep low resolution feature map after average pooling and convolution scaling through average pooling. The specific expression is and wherein represents the output of the first layer encoder, represents the output of the corresponding first layer encoder before fusion and is greater than 1, represents the output of the corresponding first layer encoder after fusion and is greater than 1, represents the cross-scale connection mapping, represents the LeakyReLU nonlinear transformation operation, represents the average pooling operation, represents the convolution operation with the kernel of 1×1×1, represents the InstanceNorm normalization operation. The encoding path comprises six encoder modules, each module comprising two AAs-conv layers, performing convolution operations with different steps to downsample the feature maps, and generating the feature maps layer by layer, and the channel number of the feature maps is increased layer by layer and the resolution is reduced layer by layer; the decoding path comprises five decoder modules, each module comprising an AAs-conv layer, performing transposed convolution to restore the resolution of the feature maps, and splicing the feature maps with the feature maps of the corresponding layer of the encoder; 2. The method of claim 1, wherein, The optimized three-dimensional medical image segmentation network model is applied to the test set to generate a three-dimensional segmentation result with clear boundaries and complete details. The preprocessing comprises the following steps: CT value truncation: the CT value of an abdominal CT image containing a kidney and a kidney tumor fine annotation is truncated in the range of [-75, 293] Hu, wherein -75 and 293 are the 0.5% quantile and the 99.5% quantile of the foreground region CT value respectively; Resampling: according to the median of the data set voxel spacing, the CT image is resampled to [3.22mm, 1.62mm, 1.62mm], the voxel spacing of each case image is unified, and the input voxel block is standardized; 3. The method of claim 1, wherein, Standardization processing: the resampled image data is unified to a standard distribution with a mean of 0 and a standard deviation of 1 by using a Z-score standardization method. The AAs-conv layer comprises the following steps: A dual-path feature extraction module: receiving an input feature map to perform dual-path parallel convolution operation, extracting two types of spatial features through a spatially separable convolution path 3x1x1-1x3x3 of a first branch and a standard three-dimensional convolution path 3x3x3 of a second branch; The attention weight generation module integrates the output features of the spatial separation convolution and the three-dimensional spatial convolution into a global feature map, compresses the spatial information of each channel of the global feature map by using average pooling, obtains the attention vectors of the branch information after convolution compression of the channel size and convolution recovery, splices the attention vectors of the branch information in the channel dimension, and then performs softmax mapping on each branch information in the channel dimension to obtain adaptive weights of the two feature maps; The weighted fusion module weights and fuses the feature maps of the two paths according to the weights of the two branches, and outputs the fused feature maps.

4. The method of claim 1, wherein, The training process includes: Input data cutting: cutting the preprocessed three-dimensional abdominal CT image into a fixed size voxel block 64X128X128 as network input, covering different sub-regions through a random center point cutting strategy; Online data enhancement: during training, online data enhancement strategy is adopted to expand training samples, including random rotation within [-30°, 30°], random size scaling of 0.7 to 1.4 times, random flipping, random brightness transformation, random contrast transformation, random gamma transformation, and various transformations, and the enhanced samples include images randomly combined by the above various transformations; Multi-stage deep supervision: at the end of each layer of the 5-layer decoder, a 1x1x1 convolution layer and a softmax activation function are used to output a segmentation prediction map of the corresponding resolution, and a composite loss function is calculated for each layer of the prediction map.

5. The method of claim 4, wherein, The loss function is: The formula is ,in Represents the cross-entropy loss function. Indicates the number of sample pixels. and These represent the predicted result and the true label at the th... i The value at each pixel The natural logarithm represents the probability of a prediction being positive, and is used to measure the confidence level of a prediction being positive. The natural logarithm represents the probability of a prediction being negative, and is used to measure the confidence level of a prediction being negative.

6. A kidney and renal tumor diagnosis and treatment assistance device comprising a memory and a processor, characterized by, The memory stores the computer program, and the processor is configured to run the computer program to perform the medical image segmentation method based on adaptive anisotropic convolution as claimed in any one of claims 1-5.

Citation Information

Patent Citations

  • Image classification method based on convolution neural network

    CN107341518A

  • Medical image segmentation device based on global information perception

    CN116258933A