An Unet multi-organ segmentation method, system, device and medium that combines adjacent layer feature fusion with large kernel convolution
Through the adjacent layer features of large nuclear convolution and the Unet multi-organ segmentation method, combined with large nuclear residual connection and GRN channel response module, the problems of high computational complexity and poor segmentation effect of multi-organ segmentation in the prior art are solved, and efficient multi-organ segmentation of medical images are achieved.
Patent Information
- Application Number
- CN202310707529.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-15
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-06-15
AI Technical Summary
The prior art has problems in the multi-organ segmentation of medical imaging, such as poor segmentation results, high computational complexity, large parameters, and poor segmentation effect on small organs in the multi-organ segmentation. In particular, the models based on convolutional neural networks have limitations in capturing remote relationships and feature fusion.
The adjacent layer feature fusion Unet multi-organ segmentation method of large-core convolution is adopted. By constructing large-core residual connections, adjacent layer feature fusions and large-core GRN channel response modules, combining large-core depth convolutions and 3×3 convolutions, the combination of local and global information is achieved, reducing computational complexity and improving segmentation performance.
It realizes efficient and efficient multi-organ segmentation of medical images under low complexity and parameter volume, reduces the complexity of organ segmentation, improves segmentation performance and efficiency, and is suitable for medical images with limited labeling samples.
Smart Images

Figure CN116681894B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-organ segmentation in medical imaging, and particularly relates to a Unet multi-organ segmentation method, system, device and medium for adjacent layer feature fusion combined with large kernel convolution. Background Technique
[0002] Organ segmentation in medical imaging is a prerequisite for basic tasks in many clinical applications such as computer-aided diagnosis (CAD) and computer-aided surgery (CAS). Automatic and accurate segmentation of multiple organs is an important but challenging task for computer-aided diagnosis and image-guided surgery systems. Accurate segmentation is an important part of many clinical applications. In current clinical practice, the contours of regions of interest are manually delineated by physicians. The process of manually delineating contours is cumbersome, time-consuming and laborious, and the organ structures and backgrounds are complex, and the organ boundaries are blurred and inconsistent. Therefore, it is challenging to manually delineate organ contours in medical images. In recent years, with the development of deep learning in the field of medical imaging. Using deep learning for multi-organ medical image segmentation can accurately draw organ contours, locate lesion areas, quickly assist doctors in locating targets, relieve the pressure on doctors, and improve the diagnostic efficiency, which has a positive and important role in modern clinical applications.
[0003] Existing technical solutions include traditional multi-organ segmentation methods, multi-organ segmentation methods based on traditional machine learning (ML), and multi-organ segmentation methods based on deep learning (DL).
[0004] Traditional segmentation methods are usually obtained from classical image segmentation methods, such as threshold-based segmentation, edge detection-based segmentation, region growing-based segmentation, etc. They involve manual operations, object boundaries, and mathematical models, and the segmentation results are poor in some cases where tissue organs are complex and the boundaries are blurred and overlapping. The multi-organ segmentation method based on traditional machine learning is driven by machine learning algorithms, such as the atlas-based segmentation method, which is a segmentation using prior knowledge. The manually segmented or pre-defined structural contours are registered to the target image through label propagation. Since the atlas-based segmentation method depends on the registration degree, and most atlases are only for the segmentation of large organs, such as the liver, kidney, spleen, chest, etc., and less for the segmentation of small organs, such as the pancreas, hepatobiliary, duodenum, esophagus, etc., the atlas-based segmentation method has deficiencies.
[0005] In recent years, deep learning technology has made great progress. Convolutional neural networks (CNNs) have been successfully applied to medical image segmentation due to their powerful feature extraction capabilities. Among different CNN variants, the U-Net network model and its variants have been advanced medical segmentation models for years due to their simple architectures and excellent performance. However, due to the inherent locality of convolutional operations, CNN-based models usually have limitations in capturing long-range relationships. Recently, with the emergence of Vision Transformer (ViT), especially the introduction of Swin Transformer (Swin-T), ViT-based models have become the backbone of medical segmentation with their superior performance. The shifted window scheme in Swin-T can overcome the limitations of high-resolution inputs while retaining the global self-attention advantages of Transformers. Although Swin-T reduces the model complexity, ViT-based models generally have large parameters, require more labeled samples and computational resources. In addition, in the field of semantic segmentation, it uses feature fusion of different layers to improve the segmentation performance, and recent works have demonstrated this, such as PSPNet and HRNet for natural image segmentation, and UNet++ and UNet 3+ for medical image segmentation. These studies show that fusing low-level features with more details and high-level features with more semantics can improve the segmentation performance in various fields. However, these methods involve fusing all features, which may lead to high computational complexity for the model, and these studies only perform simple linear mapping on the fused features, resulting in insufficient feature extraction. Summary of the Invention
[0006] In order to overcome the defects existing in the above-mentioned prior art, the purpose of the present invention is to provide a Unet multi-organ segmentation method, system, device and medium that combines adjacent layer feature fusion with large kernel convolution. By using large kernel residual connections, adjacent layer feature fusion and large kernel GRN channel responses, medical images can be effectively and efficiently segmented, achieving good segmentation performance with lower complexity and parameter quantity, and having the advantages of fewer parameters used, reduced complexity of organ segmentation, good practicability and high efficiency.
[0007] In order to achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0008] An adjacent-layer feature fusion Unet multi-organ segmentation method combined with large-kernel convolution. After dividing the dataset containing labeled samples, preprocessing the data, data sampling, and data augmentation, an adjacent-layer feature fusion UNet multi-organ segmentation model combined with large-kernel depth convolution is constructed. An adjacent-layer feature fusion module is constructed to enable the model to make full use of the information between features of different layers, obtaining lower-level features with more details and higher-level features with more semantics. A large-kernel GRN channel response module is constructed to model the long-range dependence of the features fused by adjacent layers. When the number of feature channels increases, the large-kernel GRN channel response module performs global response normalization on the channels of the fused features, thereby comparing and selecting channels, and using the fused features to improve the segmentation performance of the entire model.
[0009] The construction of the adjacent-layer feature fusion UNet multi-organ segmentation model combined with large-kernel depth convolution includes: constructing an encoder network, a basic network of the decoder, an adjacent-layer feature fusion module, and a large-kernel GRN channel response module; adding large-kernel depth convolution to the encoder network, combining the large-kernel depth convolution with 3×3 convolution to combine local information and global information.
[0010] An adjacent-layer feature fusion Unet multi-organ segmentation method combined with large-kernel convolution is as follows:
[0011] S1. Dataset division
[0012] Randomly divide the samples containing labels in the image dataset into a training set and a test set;
[0013] S2. Data preprocessing
[0014] Resample the image data after dataset division to eliminate the differences between images from different sources, facilitate the calculation and comparison of features in the images: resample the 3D CT data to the same resolution, where bilinear interpolation is used for image sample sampling and nearest interpolation is used for label sample sampling;
[0015] Normalize the data to eliminate the adverse effects caused by singular sample data and accelerate the convergence speed of network training. The formula is as follows:
[0016]
[0017] where R represents the CT data after normalization processing, Wr and Hr represent the width and height of the resolution of the CT data after normalization processing, Zr represents the number of slices, I represents the CT value before normalization processing, max(I) is the maximum CT value, and min(I) is the minimum CT value;
[0018] Centering on the target slice and stacking the upper and lower slices as the network input: First, select the target slice on the z-axis of R after normalization processing, and stack the adjacent slices centered on the target slice. Use the size of Wr×Hr×s as the network input, where s represents the number of stacked adjacent slices. If the number of stacked slices is insufficient, perform mirror padding accordingly. The whole process is as follows:
[0019] Assume that i represents a certain slice at the z-axis position in R after normalization processing. Then the network input can be expressed as:
[0020] X = [i - 1, i, i + 1]
[0021] Among them, X represents the network input, Wr and Hr represent the width and height of the original resolution of the CT data. When i = 1, that is, i is the first slice at the start of the z-axis position in R, X = [1, 1, i + 1]; when i = i max , that is, i is the last slice at the z-axis position in R, X = [i - 1, i max , i max ;
[0022] S3. Data Sampling
[0023] Sample the data X obtained from step S2 processing, that is, traverse X and perform sequential sampling; the sequential sampling method is to perform sequential sampling on X in the way of taking the slice with a step size of S, from left to right and from top to bottom, where H and W respectively represent the width and height of the slice. If the slice exceeds the size range of X during the sampling process, then sample back. If the H and W of exceed the Wr and Hr of X, then fill the surroundings with X as the center;
[0024] S4. Data Augmentation
[0025] Perform data augmentation to expand the training data and avoid overfitting caused by using few samples for training. The data augmentation realizes data augmentation in the way of horizontal flipping, vertical flipping, rotation between -90 degrees and 90 degrees, and horizontal left and right translation on the data obtained from step S3 processing according to the probability;
[0026] S5. Construct an adjacent layer feature fusion Unet multi-organ segmentation network combined with large kernel depth convolution, and name it ASF-LKUNet;
[0027] S6. Train the ASF-LKUNet segmentation network constructed in step S5;
[0028] S7. Use the test set in step S1 and the data processed in step S3 to test the optimal model trained in step S7, and quantitatively evaluate the segmentation performance of the model using the DSC coefficient and the Hausdorff distance (HD95); where HD95 calculates the distance between two sets, and the smaller the value, the smaller the distance between the two sets; its calculation method is:
[0029]
[0030]
[0031]
[0032] Among them, is the one-way Hausdorff distance between the label feature map Y and the segmentation feature map , is the one-way Hausdorff distance between the segmentation feature map and the label feature map Y, max(·) is to calculate the distance between Y and boundary points. After sorting the distances from small to large, take the sorted distances before the 95% ranking;
[0033] According to the DSC coefficient and HD95, if the DSC coefficient is higher and HD95 is lower, it means the better the segmentation performance of the model, and thus comprehensively evaluate the segmentation performance of the model.
[0034] S6. Train the ASF-LKUNet segmentation network constructed in step S5;
[0035] S7. Use the test set in step S1 and the data processed in step S3 to test the optimal model trained in step S7, and quantitatively evaluate the segmentation performance of the model using the DSC coefficient and the Hausdorff distance (HD95); where HD95 calculates the distance between two sets, and the smaller the value, the smaller the distance between the two sets; its calculation method is:
[0036]
[0037]
[0038]
[0039] Among them, is the one-way Hausdorff distance between the label feature map Y and the segmentation feature map , is the one-way Hausdorff distance between the segmentation feature map and the label feature map Y, max(·) is to calculate the distance between Y and The distance between boundary points is sorted from smallest to largest, and the sorted distances before the 95% mark are taken.
[0040] Based on the DSC coefficient and HD95, the higher the DSC coefficient and the lower the HD95, the better the model segmentation performance, and the segmentation performance of the model is comprehensively evaluated accordingly.
[0041] The specific process of data sampling in step S3 is as follows:
[0042] S301. First, calculate the padding lengths of X on the x and y axes, then we have:
[0043] P x =(S x -((H r -H) mod S x ) mod S x
[0044] P y =(S y -((W r -W) mod S y ) mod S y
[0045] Among them, P x is the padding length of X on the x-axis, P y is the padding length of X on the y-axis, S x and S y represent the step sizes in the x and y directions during sampling, S x and S y take the same step size, mod represents taking the remainder, Hr and Wr represent the width and height of the original resolution of the CT data, W and H represent the width and height of the slice If the H and W of the slice are greater than the Hr and Wr of X, then P x and P y lengths are filled on the x and y axes of X respectively. If H is greater than Hr, then P x is filled on the x-axis of X respectively. If W is greater than Wr, then P y is filled on the y-axis of X respectively;
[0046] S302. According to step S301, calculate the number of slices in the x-axis and y-axis directions;
[0047] Assume the number of slices in the x direction is N x , and its calculation method can be expressed as:
[0048] N x =(H r +P x -H) | S x +1
[0049] Among them, | represents integer division;
[0050] The number of slices N in the y direction y , and its calculation method can be expressed as:
[0051] N y =(H r +P y -H)|S y +1
[0052] S303. According to and S to calculate the coordinates, then the coordinate x' in the x direction can be expressed as:
[0053] x' = [x'1, x'2,..., x' i ,...., x' n
[0054] Among them, x' i =(i - 1)*S x , i = 1,..., n, and n represents the number of slices N in the x direction x , when i = n, x' n = Wr - W;
[0055] The coordinate y' in the y direction can be expressed as:
[0056] y' = [y'1, y'2,..., y' i ,...., y' n
[0057] Among them, y' i =(i - 1)*S y , i = 1,..., n, and n represents the number of slices N in the y direction y , when i = n, y' n = Hr - H;
[0058] S304. Use the slice step size S, and sample sequentially from left to right and from top to bottom in the xy direction of X to obtain the slice B. The specific process is as follows:
[0059] According to the coordinates of x' and y' in step S303, obtain the upper left corner coordinate v of X to be sampled:
[0060]
[0061] Among them, n = N x *N y ;
[0062] Locate on X according to the coordinate v, and intercept with the size of the slice to obtain the sampled slice B, and the formula is as follows:
[0063] B = [B1, B2,..., B n
[0064] where n = N x *N y , B i = X(v i ), B i represents the i-th slice obtained by sampling, X(v i ) represents locating on X with the coordinate v i , and then intercepting on X with the size of the slice to obtain the sampled slice.
[0065] The specific method of S5 is
[0066] S501. Construct a large kernel residual connection convolution module, including a batch normalization layer, a ReLU non-linear activation layer, a 3x3 convolution layer, and a 7x7 depth convolution layer; among them, the first 3x3 convolution layer is responsible for extracting local features and doubling the number of feature channels, the second 3x3 convolution layer is responsible for strengthening the extraction of local features, and the 7x7 depth convolution layer uses large kernel convolution to capture global features and increase the number of feature channels; apply the large kernel residual connection convolution module in the downsampling process of the model: under the operations of 3x3 and 7x7 convolutions, fuse the extracted global information and local information by addition, so that the module can capture local and global features simultaneously and effectively alleviate the limitation problem of CNN long-distance dependence modeling;
[0067] S502. Construct downsampling according to step S501. The downsampling includes a 3x3 convolution layer, a large kernel residual connection convolution module, and a max pooling layer. Among them, the 3x3 convolution layer increases the channel dimension of the initial input data of the network, the large kernel residual connection convolution module extracts features, and the max pooling layer reduces the feature resolution; every time passing through 1 max pooling layer, the feature resolution is reduced to half of the original;
[0068] S503. Construct a residual connection convolutional module, which consists of a batch normalization layer, a ReLU non-linear activation layer, and a 3x3 convolutional layer. The residual connection convolutional module first fuses the features extracted by downsampling and the upsampled features in an additive manner. The 3x3 convolutional layer with residual connection is responsible for feature extraction and reducing the number of feature channels. In the other two 3x3 convolutional layers, the first 3x3 convolutional layer is responsible for feature extraction and reducing the number of feature channels to half of the original, and the second convolution is responsible for enhancing feature extraction. The residual connection convolutional module is applied in the upsampling process of the model. The 3x3 convolutional layer is used for feature extraction, and the information extracted by the residual connection convolution and the conventional convolution is fused in an additive manner to improve the segmentation performance of the model.
[0069] S504. Construct an adjacent layer feature fusion method. During the downsampling process in step S502, adjacent layer feature fusion is performed on the extracted features. When fusing the features of three adjacent layers, it can be expressed as:
[0070]
[0071] When fusing the features of two adjacent layers, it can be expressed as:
[0072]
[0073]
[0074] Among them, and represent the fused features, represents the feature map. The subscript s represents the current scale, and the superscript (h, w, c) represents the resolution and number of channels at the corresponding scale. Conv 2×2 (·) represents a 2x2 convolutional layer with a stride of 2, and the number of output channels is twice the number of input channels. upConv 2×2 (·) represents a 2x2 transposed convolutional layer with a stride of 2, and the number of output channels is half of the number of input channels. Conv 2×2 (·) and upConv 2×2 (·) downsample and upsample the adjacent layer features to the same size as the current scale respectively;
[0075] is an operation to connect different features;
[0076] S505. Construct a GRN module, which includes three steps: global feature aggregation, feature normalization, and feature calibration;
[0077] For the feature X with an input size of (H, W, C), it can be expressed as X ∈ R H×W×C , where C is the number of feature channels, then there is:
[0078] 1) Global feature aggregation
[0079] During the global feature aggregation process, through a g function, the spatial features are aggregated into a vector, which can be expressed as:
[0080]
[0081] Among them, the above formula can obtain a value for the features of each channel by using the L2 norm, and finally obtain a set of aggregated values: In the formula, is a scalar that aggregates the statistical information of the i-th channel; assuming that X is an n-dimensional feature, that is, X = (x1, x2, x3,... x n ), then the L2 norm can be expressed as,
[0082] 2) Feature normalization
[0083] During the feature normalization process, the scalar of the statistical information of the i-th channel is normalized, which can be expressed as:
[0084]
[0085] Among them, ||X i || is the L2 norm of the i-th channel, represents the current number of channels;
[0086] 3) Feature calibration
[0087] During the feature calibration process, the original input response is calibrated using the feature normalization score calculated in step 2) of step S505, which can be expressed as:
[0088]
[0089] Among them, X i represents the i-th feature map, and respectively represent global feature aggregation and feature normalization, represents the current resolution size of X i ;
[0090] Add two additional learnable parameters γ and β, and initialize them to zero. In addition, add a residual connection between the input and output of the GRN layer. The final GRN can be expressed as:
[0091]
[0092] S506. According to step S505, construct a large kernel GRN channel response module; this module consists of layer normalization, a 7x7 depth convolution layer, a GRN layer, and a 3x3 convolution layer; among them, the 7x7 depth convolution layer is responsible for feature extraction of the features obtained by fusing the features of adjacent layers in step S504, the GRN is responsible for global response normalization of the extracted feature channels, and the 3x3 convolution layer is responsible for feature extraction and reducing the number of feature channels;
[0093] S507. Based on step S506 and step S503, construct upsampling, which consists of a residual connection convolution module, a 2x2 transposed convolution layer, and a 1x1 convolution layer; among them, the residual connection convolution module first fuses the features extracted by downsampling and the features of upsampling in an additive manner, and then extracts features. The 2x2 transposed convolution layer is responsible for increasing the feature resolution. Every time a 2x2 transposed convolution layer passes, the feature resolution increases to half of the original. The 1x1 convolution layer is responsible for mapping to the final segmentation result;
[0094] S508. According to step S502, step S504, step S506, and step S507, finally form an adjacent layer feature fusion Unet multi-organ segmentation network that combines large kernel depth convolution: ASF-LKUNet.
[0095] The specific method of step S6 is as follows:
[0096] S601. Use the cross-entropy loss function and the Dice loss function to construct the loss function of this method;
[0097] During the training of the ASF-LKUNet segmentation network, the cross-entropy loss function and the Dice loss function are defined as:
[0098]
[0099]
[0100]
[0101] where y represents the label, represents the predicted value of each category; i represents the pixel in the feature map, and c represents the category;
[0102] During the training process, the Adam (Adaptive Moment Estimation) optimizer is adopted; the partial derivative of the loss function J(θ) with respect to (θ) is calculated, the parameter θ is updated in the negative gradient direction, θ′ is the updated network parameter, and θ j is the network parameter before update, and σ is the learning rate, The training data for the input network, h θ (x i ) is the weight of the training set, y i is the label corresponding to the training set, m is the number of samples input for each training. A group of samples is randomly selected from the training set and updated according to the gradient descent rule after each training;
[0103] S602. Use the training set in the dataset constructed in step S1, the data obtained in step S3, and the data obtained by data augmentation in step S4 to train the model. During the training process, select the model with the highest evaluation index; The evaluation index uses the Dice coefficient (DSC). DSC usually measures the similarity between two samples, and its value range is [0, 1]. The higher the DSC value, the higher the similarity between the two samples. Its definition is:
[0104]
[0105] Among them, among them represents the network segmentation feature map, and Y represents the label feature map.
[0106] A segmentation system based on the above-mentioned adjacent layer feature fusion Unet multi-organ segmentation method combined with large kernel convolution includes:
[0107] A large kernel residual connection convolutional encoder for inputting network slices B and extracting global and local feature information;
[0108] A residual connection convolutional decoder for outputting a segmentation result map and extracting multi-resolution depth features;
[0109] An adjacent layer feature fusion module for fusing the features of adjacent layers, capable of obtaining lower layer features with more details and higher layer features with more semantics;
[0110] A large kernel GRN channel response module for performing global response normalization on the fused feature channels to enhance channel selection, and enhancing global and local information extraction through large kernel depth convolution, enabling the model to make full use of different features and improving the model's ability to capture global and local information.
[0111] A segmentation device based on the above-mentioned adjacent layer feature fusion Unet multi-organ segmentation method combined with large kernel convolution includes:
[0112] A memory for storing computer programs, data, and models;
[0113] A processor for implementing the operation of the landslide recognition method based on the evolutionary pruning lightweight convolutional neural network according to any one of steps 1 to 7 when executing the computer program.
[0114] A computer-readable storage medium, characterized by being responsible for reading and storing programs and data. The computer-readable storage medium stores a computer program, which, when executed by a processor, can segment organ images based on the adjacent layer feature fusion Unet multi-organ segmentation method combining large kernel convolution described in steps S1 to S7.
[0115] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0116] 1. The present invention proposes a large kernel residual connection convolution method (LK Residual Block). Among them, the residual connection can promote training, alleviate degradation and overfitting problems. Especially for medical images with limited labeled samples, large kernel depth convolution is used in the residual connection part. Combining large and small convolutions, it can capture local and global information simultaneously, effectively alleviate the limitation of CNN long-distance dependence modeling, and the large kernel depth convolution can have the ability to capture global information like ViT. Compared with the ViT architecture, it has fewer parameters, requires less labeled data and computing resources. The LK Residual Block method obtains better segmentation results compared with conventional residual connection methods, such as ResUnet.
[0117] 2. Aiming at the problems that the fully connected fusion feature method will lead to high computational complexity of the model, and the fused features cannot be effectively utilized and explored deeply enough, an adjacent layer feature fusion and large kernel global response normalization (GRN) channel response method (LKGRN) is proposed. In the adjacent layer feature fusion method, different from the fully connected fusion feature method, the adjacent layer feature fusion method fuses adjacent features in series, which can effectively reduce the computational complexity, and can fuse low-level features with more details and high-level features with more semantics, thereby improving the segmentation performance. The LKGRN method adaptively selects more meaningful channel information for the fused features and enhances feature extraction between channels through the large kernel depth convolution channel response improved based on GRN. Among them, GRN can increase the contrast and selectivity of channels, explore the relationship between channels, can effectively utilize and focus on the fused features, and will not generate additional parameters. And using large kernel depth convolution in the LKGRN method further reduces the complexity and effectively alleviates the local attention problem, and its superior performance is proved on the multi-organ dataset.
[0118] 3. The method of the present invention achieves good segmentation performance with low complexity and parameter quantity. By using large kernel residual connection, adjacent layer feature fusion and large kernel GRN channel response method, medical images can be segmented effectively and efficiently.
[0119] In summary, the present invention has the advantages of using fewer parameters, reducing the complexity of organ segmentation, good practicability and high efficiency. Brief Description of the Drawings
[0120] Figure 1 This is the flowchart of the present invention.
[0121] Figure 2 This is the overall structure diagram of ASF-LKUNet of the present invention.
[0122] Figure 3 This is the diagram of the large kernel residual connection convolutional module of the present invention.
[0123] Figure 4 This is the diagram of the residual connection convolutional module of the present invention.
[0124] Figure 5 This is the diagram of the large kernel GRN channel response module of the present invention. Detailed Embodiment
[0125] The present invention will be further described in detail below with reference to the drawings and embodiments.
[0126] A multi-organ segmentation method of Unet with adjacent layer feature fusion combined with large kernel convolution. After dividing the dataset containing labeled samples, preprocessing the data, data sampling, and data augmentation, a multi-organ segmentation model of Unet with adjacent layer feature fusion combined with large kernel depth convolution is constructed, including constructing an encoder network, a basic network of the decoder, an adjacent layer feature fusion module, and a large kernel GRN channel response module; adding large kernel depth convolution to the encoder network, combining the large kernel depth convolution with 3×3 convolution to combine local information and global information; constructing an adjacent layer feature fusion module to enable the model to make full use of the information between features of different layers to obtain lower-level features with more details and higher-level features with more semantics; constructing a large kernel GRN channel response module to perform long-range dependence modeling on the features of adjacent layer feature fusion. When the number of feature channels increases, the large kernel GRN channel response module performs global response normalization on the channels of the fused features, thereby comparing and selecting the channels, and using the fused features to improve the segmentation performance of the entire model.
[0127] See Figure 1 , a multi-organ segmentation method of Unet with adjacent layer feature fusion combined with large kernel convolution, the specific steps are as follows:
[0128] S1. Dataset Division
[0129] Randomly divide the samples containing labels in the image dataset into a training set and a test set;
[0130] S2. Data Preprocessing
[0131] Resample the impact data after partitioning the dataset to eliminate differences between images from different sources, facilitate the calculation and comparison of features in the images: Resample the 3D CT data to the same resolution of 1×1×3 mm 3 , where bilinear interpolation is used for image sample resampling and nearest interpolation is used for label samples;
[0132] The absorption of X-rays by various tissue structures in the human body varies, and different organs form different CT values. This invention mainly focuses on the segmentation of human organs such as the aorta, gallbladder, left kidney, right, liver, pancreas, spleen, and stomach. Therefore, the CT value range is selected as [-125, 275]. In addition, the data is normalized to eliminate the adverse effects caused by singular sample data and accelerate the convergence speed of network training. The formula is as follows:
[0133]
[0134] where R represents the CT data after normalization processing, Wr and Hr represent the width and height of the resolution of the CT data after normalization processing, Zr represents the number of slices, I represents the CT value before normalization processing, max(I) is the maximum CT value, and min(I) is the minimum CT value; max(I) is represented as 275 and min(I) is -125.
[0135] Take the target slice as the center and stack the upper and lower slices as the network input: First, select the target slice on the z-axis of R after normalization processing, and stack adjacent slices centered on the target slice. The size of Wr×Hr×s is used as the network input, where s represents the number of adjacent slices stacked. If the number of stacked slices is not enough, mirror filling is performed accordingly. The whole process is as follows:
[0136] Assume that i represents a certain slice at the z-axis position in the 3D sample. Then the network input can be expressed as:
[0137] X = [i - 1, i, i + 1]
[0138] where X represents the network input, Wr and Hr represent the width and height of the original resolution of the CT data. When i = 1, that is, i is the first slice at the start of the z-axis position in the 3D sample, X = [1, 1, i + 1]; when i = i max , that is, i is the last slice at the z-axis position in the 3D sample, X = [i - 1, i max , i max ;
[0139] S3. Data Sampling
[0140] Since the sample resolutions of different CT data are inconsistent and the input of the network requires a fixed resolution, the size of the network input is proportional to the network training time. Therefore, in order to reduce the network training time and ensure that all content in each data sample participates in the training, the present invention uses a size of 256x256x3 as the network input and samples the data X obtained by processing step S2, that is, traverses X and sequentially samples slices of size 256x256x3; the sequential sampling method is to use the slice with a step size of S, and perform sequential sampling on X from left to right and from top to bottom, where H and W represent 256. If the slice exceeds the size range of X during the sampling process, then sample backward. If the H and W of are greater than the Wr and Hr of X, then fill around with X as the center. The specific process is as follows:
[0141] S301. First, calculate the padding lengths of X on the x and y axes, then there are:
[0142] P x =(S x -((H r -H)modS x )modS x
[0143] P y =(S y -((W r -W)modS y )modS y
[0144] where P x is the padding length of X on the x axis, P y is the padding length of X on the y axis, S x and S y represent the step sizes in the x and y directions during the sampling process. The step size S x and S y of the present invention are both 128, mod represents taking the remainder, Hr and Wr represent the width and height of the original resolution of the CT data, W and H represent the width and height of the slice . In the present invention, H and W take 256. If the H and W of are greater than the Hr and Wr of X, then fill P x and P y lengths on the x and y axes of X respectively. If H is greater than Hr, then fill P x on the x of X respectively. If W is greater than Wr, then fill P y on the y of X respectively;
[0145] S302. Calculate the number of slices in the x-axis and y-axis directions according to step S301;
[0146] Assume the number of slices in the x direction is N x , and its calculation method can be expressed as:
[0147] N x =(H r +P x -H)|S x +1
[0148] where | represents integer division;
[0149] The number of slices N in the y direction y , and its calculation method can be expressed as:
[0150] N y =(H r +P y -H)|S y +1
[0151] S303. Calculate the coordinates according to and S, then the coordinate x' in the x direction can be expressed as:
[0152] x'=[x'1,x'2,...,x' i ,....,x' n
[0153] where, x' i =(i - 1)*S x , i = 1,..., n, n represents the number of slices N in the x direction x , when i = n, x' n = Wr - W;
[0154] The coordinate y' in the y direction can be expressed as:
[0155] y'=[y'1,y'2,...,y' i ,....,y' n
[0156] where y' i =(i - 1)*S y , i = 1,..., n, n represents the number of slices N in the y direction y , when i = n, y' n = Hr - H;
[0157] S304. Use the slice step size S to sample sequentially from left to right and top to bottom in the xy direction of X, and obtain the slice B by sampling. The specific process is as follows:
[0158] Based on the coordinates of x' and y' in step S302, obtain the upper left corner coordinates v of X to be sampled:
[0159]
[0160] where n = N x *N y ;
[0161] Locate on X according to the coordinates v, and intercept with the size of the slice to obtain the sampled slice B, and its formula is as follows:
[0162] B = [B1, B2,..., B n
[0163] where n = N x *N y , Bi i = X(v i ), Bi i represents the i-th slice obtained by sampling, X(v i ) represents locating on X with the coordinates v i , and then intercepting the sampled slice on X with the size of the slice ;
[0164] S4. Data augmentation
[0165] Since the data samples are scarce, data augmentation is needed to expand the training data and avoid overfitting caused by using few samples for training. Data augmentation is implemented by horizontally flipping, vertically flipping, rotating between -90 degrees and 90 degrees, and horizontally translating left and right the data obtained in step S3 according to a probability to achieve data augmentation;
[0166] S5. Construct a multi-organ segmentation network
[0167] Construct an adjacent layer feature fusion Unet multi-organ segmentation network combined with large kernel depth convolution, and name it ASF-LKUNet. Please refer to Figure 2 ;
[0168] S501. Construct a large kernel residual connection convolution module, which is composed of 2 batch normalization layers, 2 ReLU non-linear activation layers, 2 3x3 convolution layers and 1 7x7 depth convolution layer. Please refer to Figure 3 Among them, the first 3x3 convolutional layer is responsible for extracting local features and doubling the number of feature channels. The second convolution is responsible for strengthening the extraction of local features. The 7x7 depth convolutional layer uses large kernel convolution to capture global features and increase the number of feature channels. This module is applied in the downsampling process of the model. Under the operations of 3x3 and 7x7 different convolutions, the extracted global information and local information are fused by addition, enabling this module to capture local and global features simultaneously, effectively alleviating the limitation problem of CNN long-distance dependence modeling. Moreover, compared with the segmentation model of the ViT architecture, the 7x7 depth convolutional layer has fewer parameters, requires less labeled data and computing resources;
[0169] S502. Construct downsampling according to S501, including 1 3x3 convolutional layer, 4 large kernel residual connection convolutional modules, and 4 2x2 max pooling layers; among them, the 3x3 convolutional layer is responsible for increasing the channel dimension of the initial input data of the network, the large kernel residual connection convolutional module is responsible for feature extraction, and the max pooling layer is responsible for reducing the feature resolution. After each max pooling layer, the feature resolution is reduced to half of the original;
[0170] S503. Construct a residual connection convolutional module, which consists of 2 batch normalization layers, 2 ReLU non-linear activation layers, and 3 3x3 convolutional layers; please refer to Figure 4 ; the residual connection convolutional module first fuses the features extracted by downsampling and the features of upsampling by addition. The 3x3 convolutional layer with residual connection is responsible for feature extraction and reducing the number of feature channels. In the other two 3x3 convolutional layers, the first 3x3 convolutional layer is responsible for feature extraction and reducing the number of feature channels to half of the original, and the second convolution is responsible for strengthening feature extraction; the residual connection convolutional module is applied in the upsampling process of the model to extract features using 3 3x3 convolutional layers, and fuses the information extracted by residual connection convolution and conventional convolution by addition to improve the segmentation performance of the model;
[0171] S504. Construct an adjacent layer feature fusion method. During the downsampling process in step S502, adjacent layer feature fusion is performed on the extracted features. When fusing the features of 3 adjacent layers in a row, it can be expressed as:
[0172]
[0173] When fusing the features of 2 adjacent layers, it can be expressed as:
[0174]
[0175]
[0176] Among them, and Represents the fused features, represents the feature map. The subscript s represents the current scale, and the superscript (h, w, c) represents the resolution and number of channels at the corresponding scale. Conv 2×2 (·) represents a 2x2 convolutional layer with a stride of 2, and the number of output channels is twice the number of input channels. upConv 2×2 (·) represents a 2x2 transposed convolutional layer with a stride of 2, and the number of output channels is half the number of input channels. Conv 2×2 (·) and upConv 2×2 (·) downsample and upsample the features of adjacent layers to the same size as the current scale respectively; is an operation to connect different features;
[0177] S505. Build the GRN module. GRN includes three steps: global feature aggregation, feature normalization, and feature calibration.
[0178] For the feature X with an input size of (H, W, C), it can be expressed as X ∈ R H×W×C , then there is:
[0179] 1) Global feature aggregation
[0180] During the global feature aggregation process, through a g function, the spatial features are aggregated into a vector, which can be expressed as:
[0181]
[0182] Among them, the above formula can obtain a value for the features of each channel by using the L2 norm, and finally obtain a set of aggregated values: In the formula, is a scalar that aggregates the statistical information of the i-th channel; assuming X is an n-dimensional feature, that is, X = (x1, x2, x3,... x n ), then the L2 norm can be expressed as,
[0183] 2) Feature normalization
[0184] During the feature normalization process, the scalar of the statistical information of the i-th channel is normalized, which can be expressed as:
[0185]
[0186] Among them, ||X i || is the L2 norm of the i-th channel, represents the current number of channels;
[0187] 3) Feature calibration
[0188] During the feature calibration process, the original input response is calibrated using the feature normalization scores calculated in step 2) of step S505, which can be expressed as:
[0189]
[0190] where X i represents the i-th feature map, and represent global feature aggregation and feature normalization respectively, represents the resolution size of the current X i ;
[0191] To simplify the optimization, two additional learnable parameters γ and β need to be added and initialized to zero. Additionally, a residual connection is added between the input and output of the GRN layer. The final GRN can be expressed as:
[0192]
[0193] S506. Based on step S505, construct a large kernel GRN channel response module. This module consists of layer normalization, a 7x7 depth convolution layer, a GRN layer, and a 3x3 convolution layer; please refer to Figure 5 . Among them, the 7x7 depth convolution layer is responsible for feature extraction of the features obtained from the adjacent layer feature fusion in step S504. The GRN is responsible for global response normalization of the extracted feature channels. The 3x3 convolution layer is responsible for feature extraction and reducing the number of feature channels. Therefore, in the large kernel convolution GRN channel response module, the features of adjacent layers are fused, combining low-level details with high-level semantics, and enhancing global and local information extraction and channel selection through the large kernel channel response improved based on GRN, so that the model can make full use of different features and improve the model's ability to capture global and local information.
[0194] S507. Based on step S506 and step S503, construct upsampling. The upsampling consists of 4 residual connection convolution modules, 4 2x2 transposed convolution layers, and 1 1x1 convolution layer. Among them, the residual connection convolution module first fuses the features extracted by downsampling and the upsampled features in an additive manner, and then extracts features. The 2x2 transposed convolution layer is responsible for increasing the feature resolution. After passing through each 2x2 transposed convolution layer, the feature resolution is increased to twice the original. The 1x1 convolution layer is responsible for mapping to the final segmentation result;
[0195] According to step S502, step S504, step S506, and step S507, finally form an Unet multi-organ segmentation network that combines adjacent layer feature fusion with large kernel depth convolution: ASF-LKUNet; please refer to Figure 2 .
[0196] S6. The ASF-LKUNet segmentation network constructed in training step S5;
[0197] S601. Construct the loss function of this method using the cross-entropy loss function and the Dice loss function;
[0198] During the training of the segmenter, the cross-entropy loss function is used as the loss function and the Dice loss function is defined as:
[0199]
[0200]
[0201]
[0202] where y represents the label, represents the predicted value for each class; i represents the pixel in the feature map, and c represents the class;
[0203] During the training process, the Adam (Adaptive Moment Estimation) optimizer is adopted; first, the partial derivative of the loss function J(θ) with respect to (θ) is calculated, and the parameter θ is updated in the negative gradient direction, θ′ is the updated network parameter, and θ j is the network parameter before update, σ is the learning rate, is the training data input to the network, h θ (x i ) is the weight of the training set, y i is the label corresponding to the training set, m is the number of samples input for each training, a group of samples is randomly selected from the training set, and it is updated according to the gradient descent rule after each training;
[0204] S602. Use the training set in the dataset constructed in step S1, the data obtained in step S3, and the data obtained by data augmentation in step S4 to train the model. During the training process, select the model with the highest evaluation index; the Dice coefficient (DSC) is used as the evaluation index. DSC usually measures the similarity between two samples, and its value ranges from [0, 1]. The higher the DSC value, the higher the similarity between the two samples. It is defined as:
[0205]
[0206] where, where represents the network segmentation feature map, and Y represents the label feature map;
[0207] S7. Use the test set in step S1 and the data processed in step S3 to test the optimal model trained in step S7, and quantitatively evaluate the segmentation performance of the model using the DSC coefficient and the Hausdorff distance (HD95); among them, HD95 calculates the distance between two sets, and the smaller the value, the higher the similarity between the two sets; its calculation method is:
[0208]
[0209]
[0210]
[0211] Among them, is the one-way Hausdorff distance between the label feature map Y and the segmentation feature map , is the one-way Hausdorff distance between the segmentation feature map and the label feature map Y, max(·) is to calculate the distance between the boundary points of Y and . After sorting the distances from small to large, take the distance ranked at 95%.
[0212] According to the DSC coefficient and HD95, if the DSC coefficient is higher and HD95 is lower, it means that the model segmentation performance is better, and thus comprehensively evaluate the segmentation performance of the model.
[0213] The present invention proposes a large kernel residual connection convolution method (LK Residual Block), in which the residual connection can promote training, alleviate degradation and alleviate the overfitting problem. Especially for medical images with limited labeled samples, large kernel depth convolution is used in the residual link part. Its combination of large and small convolutions can capture local and global information simultaneously, effectively alleviate the limitation problem of CNN's long-distance dependence modeling, and the large kernel depth convolution can have the ability to capture global information like ViT. Compared with the ViT architecture, it has fewer parameters, requires less labeled data and computing resources. The LK Residual Block method obtains better segmentation results compared with conventional residual connection methods, such as ResUnet.
[0214] The present invention also addresses the problem that the fully-connected fusion feature method can lead to high computational complexity of the model, and the fused features cannot be effectively utilized and explored deeply enough. It proposes an adjacent-layer feature fusion and large-kernel global response normalization (GRN) channel response method (LKGRN). In the adjacent-layer feature fusion method, different from the fully-connected fusion feature method, the adjacent-layer feature fusion method fuses adjacent features in series, which can effectively reduce the computational complexity and can fuse low-level features with more details and high-level features with more semantics, thereby improving the segmentation performance. The LKGRN method adaptively selects more meaningful channel information for the fused features and enhances feature extraction between channels through large-kernel depth convolution channel response improved based on GRN. Among them, GRN can increase the contrast and selectivity of channels, explore the relationship between channels, can effectively utilize and focus on the fused features, and does not generate additional parameters. And in the LKGRN method, large-kernel depth convolution is used to further reduce the complexity and effectively alleviate the local attention problem, and its superior performance is demonstrated on the multi-organ dataset.
Claims
1. A Unet multi-organ segmentation method for adjacent layer feature fusion combined with large kernel convolution, characterized in that After dividing the dataset containing labeled samples, preprocessing the data, sampling the data, and augmenting the data, a UNet multi-organ segmentation model that combines adjacent layer feature fusion with large kernel depth convolution is constructed, and an adjacent layer feature fusion module is constructed to enable the model to make full use of the information between features of different layers, obtaining lower-level features with more details and higher-level features with more semantics. A large kernel GRN channel response module is constructed to model the long-range dependence of the features fused between adjacent layers. When the number of feature channels increases, the large kernel GRN channel response module performs global response normalization on the channels of the fused features, thereby comparing and selecting channels, and using the fused features to improve the segmentation performance of the entire model. The construction of the UNet multi-organ segmentation model that combines adjacent layer feature fusion with large kernel depth convolution includes: constructing an encoder network, a basic network of the decoder, an adjacent layer feature fusion module, and a large kernel GRN channel response module. Large kernel depth convolution is added to the encoder network, and the large kernel depth convolution is combined with 3×3 convolution to combine local information and global information.
2. The method for multi-organ segmentation of Unet with adjacent layer feature fusion combined with large kernel convolution according to claim 1, wherein, Specifically, it includes the following steps: S1. Dataset division Randomly divide the samples containing labels in the image dataset into a training set and a test set. S2. Data preprocessing Resample the image data after dataset division to eliminate the differences between images from different sources, facilitate the calculation and comparison of features in the images: resample the 3D CT data to the same resolution, where bilinear interpolation is used for image sample resampling and nearest neighbor interpolation is used for label samples. Normalize the data to eliminate the adverse effects caused by singular sample data and accelerate the convergence speed of network training. The formula is as follows: wherein, R represents the CT data after normalization processing, and R ∈ R Hr×Wr×Zr , Wr and Hr represent the width and height of the resolution of the CT data after normalization processing, Zr represents the number of slices, I represents the CT value before normalization processing, max(I) is the maximum CT value, and min(I) is the minimum CT value; Take the target slice as the center and stack the upper and lower slices as the network input: First, select the target slice on the z-axis of R after normalization processing, and stack the adjacent slices with the target slice as the center. Use the size of Wr×Hr×s as the network input, where s represents the number of adjacent slices stacked. If the number of stacked slices is not enough, mirror padding is performed accordingly. The whole process is as follows: Assume that i represents a certain slice at the z-axis position in R after normalization processing, then the network input can be expressed as: X = [i - 1, i, i + 1] where X represents the input of the network, X ∈ R Hr×Wr×3 , Wr and Hr represent the width and height of the original resolution of the CT data. When i = 1, that is, when i is the first slice starting from the z-axis position in R, X = [1, 1, i + 1]; when i = i max , that is, when i is the last slice starting from the z-axis position in R, X = [i - 1, i max , i max ; S3. Data sampling Sample the data X obtained by processing in step S2, that is, traverse X and perform sequential sampling; the sequential sampling method is to use a slice with a step size of S, and perform sequential sampling on X from left to right and from top to bottom, where H and W represent the width and height of the slice respectively. If the slice exceeds the size range of X during the sampling process, then sample backward. If the H and W of are greater than Wr and Hr of X, then pad around with X as the center; S4. Data augmentation Perform data augmentation to expand the training data and avoid overfitting caused by training with few samples. Data augmentation is achieved by horizontally flipping, vertically flipping, rotating between -90 degrees and 90 degrees, and horizontally translating left and right the data obtained in step S3 according to a probability. S5. Construct a UNet multi-organ segmentation network that combines adjacent layer feature fusion with large kernel depth convolution, and name it ASF-LKUNet. S6. Train the ASF-LKUNet segmentation network constructed in step S5. S7. Use the test set in step S1 and the data processed in step S3 to test the optimal model trained in step S7, and quantitatively evaluate the segmentation performance of the model using the DSC coefficient and the Hausdorff distance (HD95); among them, HD95 is to calculate the distance between two sets, and the smaller the value, the smaller the distance between the two sets; its calculation method is: Among them, is the one-way Hausdorff distance between the label feature map Y and the segmentation feature map and is the one-way Hausdorff distance between the segmentation feature map and the label feature map Y, and max(·) is the distance calculated between Y and the boundary points. After sorting the distances from small to large, the sorted distances before 95% are taken; According to the DSC coefficient and HD95, if the DSC coefficient is higher and HD95 is lower, it means that the model segmentation performance is better, and thus comprehensively evaluate the segmentation performance of the model.
3. The adjacent layer feature fusion Unet multi-organ segmentation method combined with large kernel convolution according to claim 2, characterized in that, The specific process of data sampling in step S3 is as follows: S301. First, calculate the padding lengths of X on the x and y axes, then there are: P x = (S x - ((H r - H) mod S x ) mod S x P y = (S y - ((W r - W) mod S y ) mod S y Among them, P x is the filling length of X on the x-axis, and P y is the filling length of X on the y-axis. S x and S y represent the step sizes in the x and y directions during the sampling process. S x and S y take the same step size. The mod table takes the remainder. Hr and Wr represent the width and height of the original resolution of the CT data, and W and H represent the width and height of the slice . If the H and W of x are greater than the Hr and Wr of X, then fill in P y and P x on the x and y axes of X respectively. If H is greater than Hr, then fill in P y on the x-axis of X respectively. If W is greater than Wr, then fill in P y on the y-axis of X respectively; S302. According to step S301, calculate the number of slices in the x-axis and y-axis directions; Assume that the number of slices in the x direction is N x , and its calculation method can be expressed as: N x = (H r + P x - H) | S x + 1 where, | represents integer division; The number of slices N in the y direction y , which can be calculated as follows: N y = (H r + P y - H) | S y + 1 S303. Calculate the coordinates according to and S, then the coordinate x' in the x direction can be expressed as: x′ = [x1′, x2′,..., x i ′,...., x n ′] where x i ′ = (i - 1)*S x , i = 1, ..., n, where n represents the number of slices N in the x direction x , when i = n, x n ′ = Wr - W; The coordinate y′ in the y direction can be expressed as: y′ = [y1′, y2′,..., y i ′,...., y n ′] where y i ′ = (i - 1)*S y , i = 1, ..., n, where n represents the number of slices N in the y direction y , when i = n, y n ′ = Hr - H; S304. Use slicing With a step size S, sample sequentially from left to right and top to bottom in the xy direction of X to obtain a slice B. The specific process is as follows: According to the coordinates of x′ and y′ in step S303, obtain the upper left corner coordinate v of X to be sampled: where n = N x *N y ; Locate on X according to the coordinate v and intercept with the size of the slice to obtain the sampled slice B. The formula is as follows: B = [B1, B2,..., B n where n = N x *N y , B i = X(v i ), B i represents the i-th slice obtained by sampling, B i ∈ R H×W×3 , X(v i ) represents locating at coordinate v on X, and then obtaining the sampled slice by intercepting X with the size of the slice i . 4. The adjacent layer feature fusion Unet multi-organ segmentation method combining large kernel convolution according to claim 2, wherein The specific method of S5 is: S501. Construct a large kernel residual connection convolutional module, including: a batch normalization layer, a ReLU non-linear activation layer, a 3x3 convolutional layer, and a 7x7 depth convolutional layer; among them, the first 3x3 convolutional layer is responsible for extracting local features and doubling the number of feature channels, the second 3x3 convolutional layer is responsible for strengthening the extraction of local features, and the 7x7 depth convolutional layer uses large kernel convolution to capture global features and increase the number of feature channels; apply the large kernel residual connection convolutional module in the downsampling process of the model: under the operations of 3x3 and 7x7 different convolutions, fuse the extracted global information and local information by adding them together, so that this module can capture local and global features at the same time, effectively alleviating the limitation problem of long-distance dependence modeling in CNN; S502. Construct downsampling according to step S501. The downsampling includes a 3x3 convolutional layer, a large kernel residual connection convolutional module, and a max pooling layer. Among them, the 3x3 convolutional layer raises the channel dimension of the initial input data of the network, the large kernel residual connection convolutional module extracts features, and the max pooling layer reduces the feature resolution; every time passing through 1 max pooling layer, the feature resolution is reduced to half of the original; S503. Construct a residual connection convolutional module, which consists of a batch normalization layer, a ReLU non-linear activation layer, and a 3x3 convolutional layer; the residual connection convolutional module first fuses the features extracted by downsampling and the features of upsampling by adding them together. The 3x3 convolutional layer of the residual connection is responsible for extracting features and reducing the number of feature channels. In the other two 3x3 convolutional layers, the first 3x3 convolutional layer is responsible for extracting features and reducing the number of feature channels to half of the original, and the second convolution is responsible for strengthening feature extraction; the residual connection convolutional module is applied in the upsampling process of the model, uses the 3x3 convolutional layer to extract features, and fuses the information extracted by the residual connection convolution and the conventional convolution by adding them together to improve the segmentation performance of the model; S504. Construct an adjacent layer feature fusion method. During the downsampling process in step S502, perform adjacent layer feature fusion on the extracted features. When fusing the features of three adjacent layers, it can be expressed as: When fusing the features of two adjacent layers, it can be expressed as: Among them, and represent the fused features. represents the feature map, where the subscript s represents the current scale, and the superscript (h, w, c) represents the resolution and number of channels at the corresponding scale; Conv 2×2 (·) represents a 2x2 convolutional layer with a stride of 2, and the number of output channels is twice the number of input channels; upConv 2×2 (·) represents a 2x2 transposed convolutional layer with a stride of 2, and the number of output channels is half the number of input channels. Conv 2×2 (·) and upConv 2×2 (·) downsample and upsample the features of adjacent layers to the same size as the current scale respectively; C(·) is an operation to concatenate different features. S505. Construct a GRN module. The GRN includes three steps: global feature aggregation, feature normalization, and feature calibration. For a feature X with an input size of (H, W, C), it can be expressed as X ∈ R H×W×C , where C is the number of feature channels, then we have: 1) Global feature aggregation During the global feature aggregation process, through a g function, the spatial features are aggregated into a vector, which can be expressed as: G(X):=X∈R H×W×C →gx∈R C Among them, the above formula can obtain a value for the features of each channel by using the L2 norm, and finally obtain a set of aggregated values: G(X) = gx = {||X1||, ||X2||,..., ||X C ||} ∈ R C , where G(X) i = ||X i || is a scalar that aggregates the statistical information of the i-th channel; assuming that X is an n-dimensional feature, that is, X = (x1, x2, x3,... x n ), then the L2 norm can be expressed as, 2) Feature normalization During the feature normalization process, normalize the scalar of the statistical information of the i-th channel, which can be expressed as: where, ||X i || is the L2 norm of the i-th channel, and R C represents the current number of channels; 3) Feature calibration During the feature calibration process, use the feature normalization score calculated in step 2) of S505 to calibrate the original input response, which can be expressed as: X i = X i * N(G(X) i ) ∈ R H×W Among them, X i represents the i-th feature map, N(·) and G(·) respectively represent global feature aggregation and feature normalization, and R H×W represents the resolution size of the current X i ; Add two additional learnable parameters γ and β, and initialize them to zero. Additionally, add a residual connection between the input and output of the GRN layer. The final GRN can be expressed as: X i = γ * X i * N(G(X) i ) + β + X i S506. According to step S505, construct a large kernel GRN channel response module. This module consists of layer normalization, a 7x7 depth convolution layer, a GRN layer, and a 3x3 convolution layer. Among them, the 7x7 depth convolution layer is responsible for feature extraction of the features obtained by adjacent layer feature fusion in step S504, the GRN is responsible for global response normalization of the extracted feature channels, and the 3x3 convolution layer is responsible for feature extraction and reducing the number of feature channels. S507. Based on step S506 and step S503, construct upsampling. The upsampling consists of a residual connection convolution module, a 2x2 transposed convolution layer, and a 1x1 convolution layer. Among them, the residual connection convolution module first fuses the features extracted by downsampling and the upsampled features in an additive manner, and then extracts features. The 2x2 transposed convolution layer is responsible for increasing the feature resolution. After passing through each 2x2 transposed convolution layer, the feature resolution is increased to twice the original. The 1x1 convolution layer is responsible for mapping to the final segmentation result. S508. According to steps S502, S504, S506, and S507, finally form an adjacent layer feature fusion Unet multi-organ segmentation network combined with large kernel depth convolution: ASF-LKUNet.
5. The adjacent layer feature fusion Unet multi-organ segmentation method combining large kernel convolution according to claim 2, characterized in that The specific method of step S6 is as follows: S601. Use the cross-entropy loss function and the Dice loss function to construct the loss function of this method. During the process of training the ASF-LKUNet segmentation network, the cross-entropy loss function is used as the loss function and the Dice loss function is defined as: where y represents the label, representing the predicted value for each class; i represents the pixel in the feature map, and c represents the class; During the training process, the Adam (Adaptive Moment Estimation) optimizer is adopted; the partial derivative of the loss function J(θ) with respect to (θ) is calculated, and the parameter θ is updated in the negative gradient direction. θ ′ represents the updated network parameter, and θ j represents the network parameter before update. σ is the learning rate. h is the training data input into the network. θ (x i ) represents the weight of the training set, and y i represents the label corresponding to the training set. m is the number of samples input for each training. A group of samples is randomly selected from the training set and updated according to the gradient descent rule after each training. S602. Use the training set in the dataset constructed in step S1, the data obtained in step S3, and the data obtained by data augmentation in step S4 to train the model. During the training process, select the model with the highest evaluation index. The evaluation index uses the Dice coefficient (DSC). The DSC usually measures the similarity between two samples, and its value range is [0,1]. The higher the DSC value, the higher the similarity between the two samples. Its definition is: Among them, among them represents the network segmentation feature map, and Y represents the label feature map.
6. A segmentation system for a Unet multi-organ segmentation method based on adjacent layer feature fusion combined with large-core convolution according to any one of claims 1 to 5, characterized in that, Include: A large kernel residual connection convolution encoder, which is used for the network to input slice B and extract global and local feature information. A residual connection convolution decoder, which is used to output the segmentation result map and extract multi-resolution depth features. An adjacent layer feature fusion module, which is used to fuse the features of adjacent layers to obtain low-level features with more details and high-level features with more semantics; A large kernel GRN channel response module, which is used to perform global response normalization on the fused feature channels to enhance channel selection, and enhance global and local information extraction through large kernel depth convolution, enabling the model to make full use of different features and improve the model's ability to capture global and local information.
7. A segmentation device for a multi-organ segmentation method of adjacent layer feature fusion Unet combined with large kernel convolution according to any one of claims 1 to 5, comprising: A memory, which is used to store computer programs, data and models; A processor, which is used to implement the operation of the landslide recognition method based on the evolutionary pruning lightweight convolutional neural network according to any one of claims 1 to 5 when executing the computer program.
8. A computer-readable storage medium, characterized in that, Responsible for reading and storing programs and data. The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it can segment organ images based on the multi-organ segmentation method of adjacent layer feature fusion Unet combined with large kernel convolution according to any one of claims 1 to 5.
Citation Information
Patent Citations
Laparoscopic image segmentation method and system and computer storage medium
CN115984289A
Grape leaf scab image segmentation method based on cross-resolution Transform model
CN116091770A