Stroke lesion edge enhancement segmentation method based on packet deep convolution and residual connection and computer readable storage medium

CN122820752APending Publication Date: 2026-09-25HUNAN UNIV OF CHINESE MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611016772.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-09
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0008]本发明要解决的技术问题是提供一种基于分组深度卷积与残差连接的脑卒中病灶边缘增强分割方法,可在显著降低参数量的前提下增强网络对病灶边界的敏感度,提高分割边界准确性,解决病灶边界模糊导致的分割不连续问题

Benefits of technology

[0039]一、本发明通过分组深度卷积提取多方向边缘响应,分组数等于输入通道数,使每组卷积独立学习对应通道特征图特定方向的边缘模式,有效捕捉模糊病灶的多方向边界特征。不同通道的卷积核在训练中自发朝向不同空间梯度方向收敛,实现数据驱动的多方向边缘并行检测。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820752A_ABST
    Figure CN122820752A_ABST
Patent Text Reader

Abstract

The application discloses a stroke lesion edge enhancement segmentation method based on grouping deep convolution and residual connection, and comprises the following steps: S1, a data preprocessing module: acquiring multi-modal MRI data, performing voxel value truncation and standardization, and dividing a training set, a verification set and a test set; S2, an encoder module: constructing a 3D U-Net encoder with a five-level down-sampling structure, and the feature channel numbers are 32, 64, 128, 256 and 320 in sequence; S3, a shallow edge enhancement module (EEM): introducing a shallow edge enhancement module in the first two high-resolution stages of the encoder, setting an input feature map as X element of R C×D×H×W , extracting a multi-directional edge response through grouping deep convolution, performing batch normalization and ReLU activation, element-wise adding the residual connection and the original input feature map, and then performing cross-channel feature fusion through 1x1x1 convolution to output an edge-enhanced feature map; S4, a decoder module: adopting a symmetrical up-sampling structure with the encoder, gradually restoring the spatial resolution of the feature map through transposed convolution, and finally restoring a labeled MRI image. The application enhances the edge segmentation precision under the premise of significantly reducing the parameter amount in view of the segmentation discontinuity problem caused by the fuzzy stroke lesion boundary.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of medical image processing and deep learning technology, and in particular to a method for enhancing and segmenting the lesion edge in multimodal MRI images of stroke based on grouped deep convolution and residual connections. Background Technology

[0002] Stroke is a major disease threatening human health, with ischemic stroke accounting for the largest proportion. Diffusion-weighted magnetic resonance imaging (DWI) and apparent diffusion coefficient (ADC) mapping are core imaging techniques for the diagnosis of acute ischemic stroke. Precise segmentation of the infarct lesion is of great value for thrombolytic therapy decisions, prognosis assessment, and clinical research.

[0003] In recent years, deep learning-based medical image segmentation methods have made significant progress. The U-Net architecture proposed by Ronneberger et al. employs an encoder-decoder structure, fusing shallow spatial details and deep semantic information through skip connections, becoming the foundational framework for medical image segmentation. Building upon this, nnU-Net achieves adaptive configuration, achieving excellent segmentation performance without the need for tedious tuning for specific datasets, and is gradually becoming a general baseline.

[0004] However, stroke lesion segmentation still faces significant challenges. One of the core issues leading to insufficient segmentation accuracy is the blurred lesion boundaries. In the acute phase of stroke, the infarcted area exhibits low contrast with surrounding normal brain tissue, especially at the lesion edges, where grayscale transitions are smooth and lack clear texture and morphological features. Existing convolutional neural networks primarily rely on standard convolution operations to extract features, but standard convolution is insensitive to edge responses in various directions, making it difficult to effectively capture blurred boundaries.

[0005] Researchers have explored various methods to enhance edge features. Traditional image gradient-based edge enhancement methods (such as the Sobel operator and Canny edge detection) use edge extraction as a preprocessing step before concatenating it with a CNN segmentation network. However, these methods have fixed edge detection operator parameters, only detecting edges in preset directions such as horizontal, vertical, and diagonal, and cannot adaptively optimize for the segmentation task. Furthermore, edge extraction and semantic segmentation are independent, and edge information cannot guide the feature learning of the segmentation network. Attention mechanisms such as the SE module adaptively recalibrate channel weights through global pooling and fully connected layers, but this module only models inter-channel dependencies, ignoring spatial edge information. CBAM improves feature representation to some extent by concatenating channel attention and spatial attention, but its spatial attention is directly generated using a 7×7 standard convolution on the original number of channels (C), resulting in a parameter count on the C×C×49 level, leading to significant computational overhead, and it is not specifically designed for multi-directional edge responses.

[0006] It is worth noting that grouped deep convolutions are commonly found in lightweight networks (such as MobileNet and ShuffleNet), but their main purpose is to accelerate the model by reducing computational cost through grouping, rather than to enhance the ability to extract spatial edge features. Furthermore, the aforementioned existing methods typically apply the embeddings evenly across multiple layers of the network without specifically analyzing the "optimal location for edge enhancement"—in reality, shallow feature maps have high spatial resolution and rich edge details, making them the optimal stage for edge enhancement, while deep feature maps have low resolution and semantic abstraction, resulting in limited edge enhancement effects.

[0007] Therefore, there is an urgent need for a lightweight and efficient edge enhancement segmentation method that can effectively enhance the response of lesion edges in the shallow stage of the network, without significantly increasing the number of network parameters and computational burden, thereby improving the accuracy of segmentation boundaries. Summary of the Invention

[0008] The technical problem to be solved by the present invention is to provide a stroke lesion edge enhancement segmentation method based on grouped deep convolution and residual connections, which can enhance the network's sensitivity to lesion boundaries and improve the accuracy of segmentation boundaries while significantly reducing the number of parameters, thus solving the problem of segmentation discontinuity caused by blurred lesion boundaries.

[0009] To solve the above problems, the technical solution of the present invention is as follows:

[0010] A method for enhancing and segmenting the edges of stroke lesions based on grouped depthwise convolution and residual connections, the method comprising the following steps:

[0011] S1, Data Preprocessing: Acquire multimodal MRI data, filter irrelevant regions according to the voxel truncation method, standardize the filtered multimodal MRI data, and divide it into training set, validation set and test set.

[0012] S2, Component Encoder: A five-level downsampling 3D U-Net encoder is constructed. Each level contains two 3×3×3 convolutional layers, followed by instance normalization and LeakyReLU activation. Downsampling is achieved through 3×3×3 convolutions with a stride of 2. The number of feature channels is 32, 64, 128, 256, and 320 respectively. "Downsampling" refers to the operation of progressively reducing the spatial resolution of the feature map, which is one of the core functions of the encoder. In the 3D U-Net encoder used in this invention, the depth (D), height (H), and width (W) of the feature map are halved after each downsampling level. For example, if the current feature map spatial size is 64×64×64, it becomes 32×32×32 after downsampling.

[0013] S3, Shallow Edge Enhancement: A shallow edge enhancement module (EEM) is introduced in the first two high-resolution stages of the encoder. This module extracts multi-directional edge responses through grouped deep convolutions and performs residual fusion with the original features, enhancing the network's sensitivity to lesion boundaries. Specifically, the "original features" refer to the input feature map X of the shallow edge enhancement module in step S31 below.

[0014] S4, Decoding and Restoration: The decoder, which has a structure symmetrical to the encoder, gradually restores the spatial resolution of the feature map through transposed convolution, and fuses the features of the corresponding layers of the encoder through skip connections to restore the labeled MRI image.

[0015] The training process employs a size-weighted Dice-CE loss function, which consists of a weighted Dice loss and a weighted cross-entropy loss, with differentiated weights assigned based on the number of voxels in each independent lesion region.

[0016] Furthermore, the shallow edge enhancement described in step S3 specifically includes:

[0017] S31. Let the input feature map be X∈R C×D×H×W Where C is the number of channels, and D, H, and W are the depth, height, and width, respectively.

[0018] S32 extracts spatial edge features through 3×3×3 grouped depthwise convolutions, where the number of groups equals the number of input channels C, padded with 1s. Each group of convolutions independently learns the spatial edge pattern for its corresponding channel. Because each group of convolutional kernels learns independently, the kernels of different channels spontaneously converge towards different spatial gradient directions (such as horizontal, vertical, diagonal, etc.), thus achieving parallel detection of edges in multiple directions. The fundamental difference between this mechanism and traditional fixed-direction operators like Sobel is that the direction pattern is adaptively learned through data, rather than being manually preset.

[0019] S33, batch normalization (BN) and ReLU activation are sequentially applied to the outputs of the grouped depthwise convolutions to accelerate convergence and introduce nonlinearity. Batch normalization is a training optimization technique in deep neural networks, proposed by Sergey Ioffe and Christian Szegedy in 2015. Its core idea is to normalize the feature distribution of the current batch of data in each layer of the network, making its mean 0 and variance 1, thereby solving the problem of internal covariate shift in neural network training. In the EEM module of this invention, batch normalization is applied after the outputs of the grouped depthwise convolutions and before the ReLU activation function.

[0020] S34 performs feature fusion by adding the output of S33 to the original input feature map X element by element through residual connection, so as to preserve the original feature information and alleviate gradient vanishing.

[0021] S35 uses 1×1×1 convolution to achieve cross-channel feature fusion, integrating and combining the spatial edge features extracted independently from each group, and outputting an enhanced feature map X_out of the same size. The overall calculation method is as follows:

[0022] ,

[0023] Among them, Conv 3×3×3 `group` represents a 3×3×3 depthwise convolution with C groups, and `BN` represents batch normalization. 1×1×1 This represents a 1×1×1 convolution.

[0024] In the above formula, the output of the grouped depthwise convolution can be further expressed as:

[0025] Y c =X c *K c c = 1, 2, ..., C

[0026] Among them, X c Let K be the input feature map of the c-th channel. c This is a dedicated 3×3×3 convolution kernel for this channel; * indicates a three-dimensional convolution operation. Since each convolution kernel has K... c Independent learning allows convolutional kernels from different channels to spontaneously converge toward different spatial gradient directions during training, thus enabling parallel detection of edges in multiple directions.

[0027] The role of "batch normalization" in this invention is threefold: First, it accelerates network convergence. Batch normalization effectively alleviates internal covariate bias, allowing the network to safely adopt a larger learning rate and significantly reducing the number of training iterations. In the ablation experiments of this patent, the training convergence speed decreased by approximately 30% without Batch Normalization. Second, it stabilizes the training process. Batch Normalization alleviates the gradient vanishing / exploding problem in deep networks. In the EEM module, the feature distribution after Batch Normalization is controlled within a reasonable range, enabling stable gradient backpropagation, which is particularly beneficial for updating deep network parameters. Third, it provides a slight regularization effect. Due to the slight fluctuations in the mean and variance of each batch, Batch Normalization introduces a certain amount of random noise into the training process, playing a regularization role similar to Dropout, which helps improve the model's generalization ability.

[0028] Furthermore, in the shallow edge enhancement module, the kernel size of the grouped depthwise convolution is 3×3×3 with padding of 1, and the number of groups equals the number of input channels C, allowing each channel to perform spatial convolution independently. The number of parameters is 1 / C of that of a conventional 3×3×3 convolution. The residual connection adds the output of the grouped depthwise convolution path after batch normalization and ReLU activation element-wise to the original input feature map X, and then performs cross-channel feature fusion through 1×1×1 convolution. This ensures that the output feature map retains the original feature information while incorporating multi-directional edge enhancement features. This module reduces the number of parameters to 1 / C of a standard 3×3×3 convolution through grouped depthwise convolution, with a total number of parameters of approximately C×C, which is significantly reduced compared to C2×27 of a conventional 3×3×3 convolution.

[0029] Furthermore, in the shallow edge enhancement module, the grouped deep convolution structure differs fundamentally from existing grouped convolution applications: existing lightweight networks (such as MobileNet) primarily use grouped convolution to split computations along the channel dimension to reduce overall computational cost; the purpose of this grouping operation is model acceleration, not specific extraction of edge features. This invention sets the number of groups in the grouped deep convolution to be equal to the number of input channels C, allowing each channel to undergo independent spatial convolution. Each group of convolutions independently learns the local spatial gradient pattern of the channel's feature map, thereby extracting multi-directional spatial edge responses. This design differs from general attention mechanisms such as CBAM in that: CBAM's spatial attention module directly generates a spatial attention map using a 7×7 standard convolution on the original number of channels C, with a parameter count of C×C×49, and does not distinguish the edge direction characteristics of each channel; this invention first extracts the edge features of each channel independently through grouped deep convolution (parameter count C×3×3×3=C×27), and then fuses multi-directional edge responses through a 1×1×1 convolution (parameter count C×C), with a total parameter count of approximately C×27+C. 2 It has a higher level of computational efficiency, and the direction of edge detection is driven by data-driven adaptive learning rather than being preset and fixed.

[0030] Furthermore, the shallow edge enhancement module only operates on the first two high-resolution stages of the encoder, namely the stages with 32 and 64 channels, and does not operate on the deep low-resolution stages. The output of the shallow edge enhancement module is an enhanced feature map of the same size, which continues to be passed down to the next encoder level. Its design is based on the fact that shallow feature maps have high spatial resolution and rich edge details, making them the optimal location for edge enhancement; deep feature maps have low spatial resolution, large receptive fields, and abstract semantic information, limiting the effectiveness of edge enhancement at this stage. By applying edge enhancement in the shallow layers, the network can more accurately capture the boundary transition between lesions and normal tissue, providing clearer boundary cues for subsequent decoders.

[0031] Furthermore, in the grouped depthwise convolution, the convolutional kernels of different channels spontaneously converge toward different spatial gradient directions during training, including horizontal, vertical and diagonal directions, thereby achieving parallel detection of multi-directional edges; the direction pattern is driven by data-driven adaptive learning rather than manually preset fixed directions.

[0032] Furthermore, the input features received by the decoder include: the decoder's own features after upsampling, and the skip connection features from the corresponding level of the encoder. The first two skip connection features are enhanced by a shallow edge enhancement module, enabling the decoder to utilize the enhanced edge information for more accurate boundary localization when restoring spatial resolution. The decoder output is mapped to the number of target categories via a 1×1×1 convolution, and a probability map of each voxel belonging to a lesion is generated by a Sigmoid activation function. A threshold of 0.5 is then applied to obtain the final segmentation result. Upsampling is an operation in deep learning that progressively increases the spatial resolution of the feature map; it is the inverse of downsampling. In the 3D U-Net decoder, the goal of upsampling is to restore the feature map compressed during the encoder stage to the spatial size of the original input image, thereby achieving pixel-by-pixel classification prediction.

[0033] Furthermore, the infrastructure of the encoder and decoder follows the adaptive configuration framework of nnU-Net, including automatic data resampling, intensity normalization, and patch size and batch size configuration based on GPU memory limitations.

[0034] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.

[0035] MRI, short for Magnetic Resonance Imaging, is an advanced medical imaging technique that uses strong magnetic fields and radio frequency pulses to cause hydrogen nuclei within human tissues to resonate. The emitted signals are then received and used by a computer to reconstruct high-resolution images of the body's internal structures. Unlike CT scans, MRI does not emit ionizing radiation and possesses extremely high contrast resolution for soft tissues (such as the brain, spinal cord, articular cartilage, and muscles), clearly distinguishing between gray matter, white matter, and minute lesions. Due to these characteristics, MRI plays a central role in the diagnosis and lesion segmentation research of stroke, tumors, and neurodegenerative diseases.

[0036] ReLU stands for Rectified Linear Unit, and it's one of the most commonly used activation functions in deep neural networks. Its function is to filter out irrelevant negative activation information during the generation of the spatial attention weight map, allowing the network to better focus on the small lesion areas it's "interested in" and suppress background noise. In short, ReLU acts like a signal filter: it only allows useful, strong feature signals to pass through, blocking weak negative signals and noise, helping the network learn clearer and more discriminative image features.

[0037] LeakyReLU, short for Leaky Rectified Linear Unit, is a commonly used activation function in deep neural networks. It's an improved version of the standard ReLU (Rectified Linear Unit). Simply put, ReLU outputs the original value for positive inputs and 0 for negative inputs; while LeakyReLU outputs the original value for positive inputs and a small negative value (instead of 0) for negative inputs, thus preserving some negative information and solving the "neuron death" problem of ReLU.

[0038] This invention provides a method for enhancing and segmenting the edges of stroke lesions based on grouped depthwise convolution and residual connections, the advantages of which are:

[0039] I. This invention extracts multi-directional edge responses through grouped depthwise convolution, with the number of groups equal to the number of input channels. This allows each group of convolutions to independently learn edge patterns in a specific direction corresponding to the feature map of that channel, effectively capturing multi-directional boundary features of blurred lesions. During training, the convolutional kernels of different channels spontaneously converge towards different spatial gradient directions, achieving data-driven parallel detection of multi-directional edges.

[0040] Second, the module preserves the original feature information through residual connections to avoid feature degradation during the edge enhancement process. At the same time, 1×1×1 convolution realizes cross-channel information fusion, integrating the edge responses of each independent channel into a unified edge enhancement representation, thus balancing edge enhancement and semantic information preservation.

[0041] Third, grouped depthwise convolution reduces the number of parameters to 1 / C of that of standard convolution, with a total number of parameters of approximately C×C. Compared to the standard 3×3×3 convolution's C×C×27, the module introduces very little additional computational overhead and can be plugged and played into existing segmentation networks without changing the main structure of the network.

[0042] Fourth, this module is specifically deployed in the shallow high-resolution stage of the encoder, making full use of the rich edge information of the shallow feature map. While maximizing the edge enhancement effect, it avoids unnecessary computational overhead in the deep stage, thus achieving accurate and efficient edge enhancement. Attached Figure Description

[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0044] Figure 1 This is a flowchart of the steps of the stroke lesion edge enhancement segmentation method based on grouped depth convolution and residual connection of the present invention;

[0045] Figure 2 This is a schematic diagram of the network structure and shallow edge enhancement module of the present invention.

[0046] Specific implementation methods

[0047] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, and to make the above-mentioned objectives, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be further described below.

[0048] The endpoints and any values ​​of the ranges disclosed herein are not limited to the precise ranges or values, and these ranges or values ​​should be understood to include values ​​close to these ranges or values. For numerical ranges, the endpoint values ​​of the various ranges, the endpoint values ​​of the various ranges and individual point values, and individual point values ​​can be combined with each other to obtain one or more new numerical ranges, which should be considered as specifically disclosed herein.

[0049] Example 1

[0050] refer to Figure 1 and Figure 2 This embodiment presents a method for enhancing and segmenting the edges of stroke lesions based on grouped depthwise convolution and residual connections. The method includes the following steps:

[0051] S1, Data Preprocessing: Acquire multimodal MRI data, filter irrelevant regions according to the voxel truncation method, standardize the filtered multimodal MRI data, and divide it into training set, validation set and test set.

[0052] As a preferred implementation, in step S1, we acquired the ISLES-2022 (Ischemic Stroke Lesion Segmentation) dataset for segmenting lesions in acute ischemic stroke. This dataset contains multimodal MRI data of the brains of 250 patients with acute ischemic stroke, with each data point providing three raw modalities: ADC (apparent diffusion coefficient), DWI (diffusion-weighted imaging), and FLAIR (fluid attenuation inversion recovery). The volume of the data images ranges from [64-128] × [64-128] × [32-64] voxels, and the voxel spacing ranges from [1.0-2.5] mm. The infarct areas in the dataset were precisely annotated by radiologists, and the annotations contain a large number of tiny lesions with diameters of only a few millimeters. Voxel values ​​outside the 0.5%-99.5th percentile were truncated for each modality to eliminate extreme noise. Then, z-score normalization was performed to make the mean 0 and the variance 1 for each modality. The standardized parameters are calculated based on the statistical values ​​of the foreground voxels in the training set. Finally, the dataset is divided into training and validation sets using a five-fold cross-validation method, with a separate test set reserved. The validation set is used to select the model's hyperparameters.

[0053] S2, Component Encoder: A five-level downsampling 3D U-Net encoder is constructed. Each level contains two 3×3×3 convolutional layers, followed by instance normalization and LeakyReLU activation. Downsampling is achieved through 3×3×3 convolutions with a stride of 2. The number of feature channels is 32, 64, 128, 256, and 320 respectively, to accommodate the need for feature abstraction from shallow to deep. The first-level encoder receives preprocessed multimodal MRI images with an initial number of 32 feature channels. With each downsampling pass, the number of channels doubles, and the spatial resolution is halved. After five downsampling passes, the spatial size of the final-level feature map is reduced to 1 / 32 of its original size, and the number of channels increases to 320, resulting in the largest receptive field and the richest semantic information.

[0054] S3, Shallow Edge Enhancement: A shallow edge enhancement module is introduced in the first two high-resolution stages of the encoder. Multi-directional edge responses are extracted through grouped deep convolutions and residual fusion is performed with the original features to enhance the network's sensitivity to lesion boundaries.

[0055] S4, Decoding and Restoration: The decoder, which has a structure symmetrical to the encoder, gradually restores the spatial resolution of the feature map through transposed convolution, and fuses the features of the corresponding layers of the encoder through skip connections to restore the labeled MRI image.

[0056] Preferably, the shallow edge enhancement in step S3 specifically includes:

[0057] S31. Let the input feature map be X∈R C×D×H×WWhere C is the number of channels, and D, H, and W are the depth, height, and width, respectively.

[0058] S32 extracts spatial edge features through 3×3×3 grouped depthwise convolutions. The number of groups equals the number of input channels C, with padding of 1. Each group of convolutions independently learns the spatial edge pattern of its corresponding channel. The grouped depthwise convolutions divide the 64 channels into 64 independent groups, each containing only one channel, and use a 3×3×3 convolution kernel to spatially convolve that channel. This design allows each group of convolutions to independently learn the local spatial gradient pattern of its corresponding feature map, thereby extracting multi-directional spatial edge responses. The grouped depthwise convolutions have 64×1×3×3×3=1,728 parameters, while the regular 3×3×3 convolutions have 64×64×3×3×3=110,592 parameters, only 1 / 64th the number of parameters of a regular convolution.

[0059] S33 performs batch normalization and ReLU activation on the output of the grouped depthwise convolution in sequence to accelerate convergence and introduce nonlinearity.

[0060] S34, through residual connections, adds the output of S33 to the original input feature map X element-wise to perform feature fusion, thereby preserving the original feature information and mitigating gradient vanishing. The design of residual connections allows the module to retain the original feature information while enhancing edges, avoiding the loss of original useful information due to edge enhancement operations, and also helps to alleviate the gradient vanishing problem in deep network training.

[0061] S35 uses 1×1×1 convolution to achieve cross-channel feature fusion, integrating and combining the spatial edge features extracted independently from each group, and outputting an enhanced feature map X_out of the same size. The overall calculation method is as follows:

[0062] ,

[0063] Among them, Conv 3×3×3 `group` represents a 3×3×3 depthwise convolution with C groups, and `BN` represents batch normalization. 1×1×1 This represents a 1×1×1 convolution. The number of parameters in a 1×1×1 convolution is 64×64×1×1×1=4,096. The total number of parameters in the module is approximately 1,728+4,096=5,824, which is about 94.7% less than the 110,592 parameters of a regular 3×3×3 convolution.

[0064] The structure of the EEM module with 32 channels is the same as above, except that the number of channels changes from 64 to 32. The number of parameters for grouped depthwise convolution in this stage is 32×1×3×3×3=864, and the number of parameters for 1×1×1 convolution is 32×32×1×1×1=1,024, for a total of approximately 1,888 parameters.

[0065] Preferably, in the shallow edge enhancement module, the kernel size of the grouped depth convolution is 3×3×3 with padding of 1, and the number of groups is equal to the number of input channels C, so that each channel performs spatial convolution independently, and the number of parameters is 1 / C of the conventional 3×3×3 convolution; the residual connection adds the output of the grouped depth convolution path after batch normalization and ReLU activation to the original input feature map X element by element, and then completes cross-channel feature fusion through 1×1×1 convolution, so that the output feature map retains the original feature information and fuses multi-directional edge enhancement features.

[0066] Preferably, the shallow edge enhancement module only operates on the first two high-resolution stages of the encoder, i.e., the stages with 32 and 64 channels, and does not operate on the deep low-resolution stages. The output of the shallow edge enhancement module is an enhanced feature map of the same size, which is then passed down to the next encoder layer. Its design is based on the fact that shallow feature maps have high spatial resolution, preserving rich spatial details and edge information, making them the optimal location for edge enhancement. As the network deepens, the spatial resolution of the feature maps gradually decreases, while the receptive field gradually increases. Deep feature maps contain more abstract semantic information than specific edge details, resulting in limited edge enhancement effects. Limiting the EEM module to the shallow layers maximizes the edge enhancement effect while avoiding unnecessary computation in deeper layers. The module output is an enhanced feature map of the same size, which is then passed down to the next encoder layer.

[0067] Preferably, in grouped depthwise convolution, the convolution kernels of different channels spontaneously converge toward different spatial gradient directions during training, including horizontal, vertical and diagonal directions, thereby achieving parallel detection of multi-directional edges; the direction pattern is driven by data-driven adaptive learning rather than manually preset fixed directions.

[0068] Preferably, the decoder described in step S4 employs an upsampling structure symmetrical to the encoder, consisting of five stages. Each stage of the decoder first doubles the feature map size and halves the number of channels through transposed convolution, then refines the features through two 3×3×3 convolutional layers. Each convolution is followed by instance normalization and LeakyReLU activation. After five stages of upsampling, the feature map is restored to its original input resolution.

[0069] The input features received by the decoder include: the decoder's own features after upsampling, and the skip connection features from the corresponding level of the encoder. The first two skip connection features are features enhanced by the shallow edge enhancement module, which enables the decoder to use the enhanced edge information for more accurate boundary localization when restoring spatial resolution. The decoder output is mapped to the number of target categories by a 1×1×1 convolution, and a probability map of each voxel belonging to the lesion is generated by the Sigmoid activation function. The final segmentation result is obtained by taking a threshold of 0.5.

[0070] Preferably, the infrastructure of the encoder and decoder follows the adaptive configuration framework of nnU-Net, including automatic data resampling, intensity normalization, and patch size and batch size configuration based on GPU memory limitations.

[0071] This embodiment also discloses a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above method.

[0072] Preferably, the size-weighted Dice-CE loss function used in the training process of this embodiment consists of weighted Dice loss and weighted cross-entropy loss, and is defined as follows:

[0073] ,

[0074] Among them, w ce =1.5, w dice =0.8, w ce and w dice All are balance coefficients;

[0075] The size weighting method is as follows: Three-dimensional connected component analysis is performed on the independent lesion regions in each training sample, the voxel count v of each lesion is counted, and differential weights are assigned based on the volume distribution.

[0076] ,

[0077] This allows voxels with small lesions to receive a higher loss weight and voxels with large lesions to receive a lower loss weight, thus mitigating the class imbalance problem.

[0078] Preferably, the weighted Dice loss is defined as:

[0079] ;

[0080] The weighted cross-entropy loss is defined as:

[0081] ;

[0082] Among them, w p Let p be the size weight. Predict the probability that position p belongs to category c for the network. Let C be the true label at position p, C be the number of categories, and ε be the smoothing term with a value of 1 × 10⁻ 5 To prevent the denominator from being zero, the size weight map is normalized before calculating the weighted loss, so that the average weight of the entire map is 1, in order to keep the magnitude of the loss function consistent with that of the unweighted map.

[0083] The method of this invention was successfully applied to the ISLES-2022 dataset and achieved segmentation, employing a five-fold cross-validation strategy. To verify the segmentation performance of the method of this invention on lesions of different sizes, the lesions in the test set were divided into four levels according to the number of voxels: micro lesions (v<50), small lesions (50≤v<200), medium lesions (200≤v<800), and large lesions (v≥800). The Dice coefficients at each scale were calculated, and the results are shown in Table 1.

[0084] Table 1 Comparison of lesion segmentation performance at different voxel scales

[0085]

[0086] Note: Dice = Dice coefficient based on case average (%), HD95 = 95% Hausdorff distance (mm), AVD = absolute volume difference, LDR = lesion detection rate (%), MMC = volume Pearson correlation coefficient.

[0087] As shown in Table 1, the method of the present invention achieves good segmentation results on lesions of different sizes. The Dice coefficient for small lesions reaches 77.94%, and for large lesions it reaches 83.12%, proving that the method has robust segmentation performance for lesions of different sizes.

[0088] To further verify the effectiveness of the shallow edge enhancement module, ablation experiments were conducted. The shallow edge enhancement module (EEM) was added to the baseline nnU-Net. The results are shown in Tables 2-1 and 2-2.

[0089] Table 2-1 Ablation Experiment Results (Overlap and Distance Indicators)

[0090]

[0091] Note: Dice_case = Dice coefficient (%) based on case average, Dice_global = Dice coefficient (%) based on global pixels, HD95 = 95% Hausdorff distance (mm), MSD = average surface distance (mm), F1_2mm = F1 (%) based on points with a 2mm tolerance, F1_5mm = F1 (%) based on points with a 5mm tolerance. ↑ indicates a larger value for better performance, ↓ indicates a smaller value for better performance.

[0092] Table 2-2 Ablation test results (volume and detection index)

[0093]

[0094] Note: AVD = Absolute Volume Difference, Detection Rate = Lesion Detection Rate, Volume_Pearson_r = Volume Pearson Correlation Coefficient. ↑ indicates a larger value indicates better performance, ↓ indicates a smaller value indicates better performance.

[0095] As shown in Tables 2-1 and 2-2, after adding the EEM module to the baseline nnU-Net, the Dice coefficient based on case average increased from 72.69% to 72.33%, remaining essentially unchanged; HD95 decreased significantly from 5.8705 mm to 4.8788 mm, a decrease of 16.9%; MSD decreased from 2.3498 mm to 2.2176 mm, a decrease of 5.6%; lesion detection rate increased from 37.82% to 40.16%, an increase of 2.34 percentage points; and the volumetric Pearson correlation coefficient increased from 0.1903 to 0.2183. These improvements validate that the EEM module, by applying grouped deep convolutions and residual connections in the shallow stages of the encoder, effectively enhances the network's sensitivity to lesion boundaries, significantly improving boundary segmentation accuracy and lesion detection rate while significantly reducing the number of parameters.

[0096] To further verify the superiority of the shallow edge enhancement module over the classic attention mechanism, a comparative experiment was conducted between the EEM module of this invention and the convolutional block attention module (CBAM). The results are shown in Tables 3-1 and 3-2.

[0097] Table 3-2 Comparison of EEM and CBAM experimental results (volume and detection index)

[0098]

[0099] Note: The meanings of the indicators are the same as in Table 2-2.

[0100] As shown in Tables 3-1 and 3-2, the EEM module outperforms or matches CBAM in several metrics. Regarding boundary accuracy, EEM's HD95 is 7.4804 mm, better than CBAM's 8.0832 mm, a decrease of 7.5%; its MSD is 2.9581 mm, better than CBAM's 3.7561 mm, a decrease of 21.2%. In terms of boundary fineness, EEM's F1_2 mm and F1_5 mm reach 28.31% and 33.06% respectively, both better than CBAM's 26.31% and 30.76%. In the Dice_global metric, EEM slightly outperforms CBAM's 78.60% with 78.73%. These results demonstrate that the EEM module's design, which extracts multi-directional edge responses through grouped depthwise convolution, has significant advantages in boundary localization accuracy and segmentation consistency compared to CBAM's method of directly generating spatial attention maps using 7×7 standard convolutions. Furthermore, it requires fewer parameters and has higher computational efficiency.

[0101] This embodiment also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.

[0102] Example 2

[0103] The difference between this embodiment and Embodiment 1 is that the kernel size of the grouped depthwise convolution in the shallow edge enhancement module is 5×5×5 instead of 3×3×3. The number of groups is still equal to the number of input channels C, and the padding is 2. Using a larger convolution kernel can capture a wider range of edge context information, but the number of parameters increases accordingly (from C×27 to C×125). Taking a stage with 64 channels as an example, the number of parameters for the 5×5×5 grouped depthwise convolution is 64×1×5×5×5=8,000, and the number of parameters for the 1×1×1 convolution is 4,096, for a total of approximately 12,096 parameters. Experiments on the ISLES-2022 dataset show that the EEM module with the 5×5×5 convolution kernel is slightly better than the 3×3×3 version in terms of boundary accuracy (HD95) (4.72mm vs 4.88mm), but the number of parameters increases by about 2 times, and the training speed decreases by about 15%. Considering both accuracy and efficiency, 3×3×3 is the preferred solution.

[0104] Example 3

[0105] The difference between this embodiment and Embodiment 1 is that the shallow edge enhancement module does not use batch normalization layers, but only retains grouped depthwise convolutions, ReLU activations, residual connections, and 1×1×1 convolutions. Experimental results show that the training convergence speed is reduced by about 30% without batch normalization, and the final HD95 index is 5.12mm, slightly worse than the 4.88mm with batch normalization, verifying the role of batch normalization in accelerating convergence and improving training stability.

[0106] Example 4

[0107] The difference between this embodiment and Embodiment 1 is that residual connections are not used in the shallow edge enhancement module. That is, the output of the grouped depthwise convolution is directly input into a 1×1×1 convolution after batch normalization and ReLU activation, without being added to the original input feature map X. Experimental results show that without residual connections, some of the original spatial information is lost in the module's output feature map, and HD95 increases from 4.88mm to 5.34mm, verifying the importance of residual connections in preserving original feature information.

[0108] Example 5

[0109] The difference between this embodiment and Embodiment 1 is that the encoder's basic structure adopts the standard 3D U-Net instead of the nnU-Net adaptive configuration framework. In the standard 3D U-Net, the number of feature channels in the encoder's five downsampling levels are 32, 64, 128, 256, and 512, respectively. The shallow edge enhancement module still operates in the first two stages with 32 and 64 channels. This embodiment verifies the universality of the method of the present invention under different encoder architectures. The results show that after adding the EEM module to the standard 3D U-Net architecture, HD95 is reduced from 6.8mm to 5.2mm, verifying the effectiveness of the EEM module of the present invention on different backbone networks.

[0110] Comparative Example 1

[0111] The difference between this comparative example and Example 1 is that a standard 3×3×3 convolution is used instead of grouped depthwise convolution to extract edge features; that is, the grouping strategy is not adopted, while the rest of the structure remains unchanged. The standard convolution has C×C×27 parameters, which is 110,592 parameters for 64 channels, while the grouped depthwise convolution has only 1,728 parameters. Experimental results show that the HD95 of the standard convolution version is 4.95mm, slightly worse than the 4.88mm of the grouped depthwise convolution, and the number of parameters increases by 64 times, verifying the effectiveness of the grouped depthwise convolution in maintaining or even improving edge extraction capabilities while reducing the number of parameters.

[0112] Comparative Example 2

[0113] The difference between this comparative example and Example 1 is that the shallow edge enhancement module is deployed in all five stages of the encoder (EEM is applied to channels 32, 64, 128, 256, and 320), rather than just the first two high-resolution stages. Experimental results show that applying EEM in the deeper stages not only fails to significantly improve boundary accuracy (4.92mm vs 4.88mm for HD95), but also increases computational overhead by approximately 15%, validating the rationality of limiting EEM to the shallow stages.

[0114] Comparative Example 3

[0115] The difference between this comparative example and Example 1 is that no edge enhancement module is introduced; the baseline nnU-Net is used directly for segmentation. Experimental results show that the HD95 of the baseline model is 5.87 mm, which is significantly worse than the 4.88 mm after adding EEM (a decrease of 16.9%), verifying the core role of the EEM module in improving boundary segmentation accuracy.

[0116] The embodiments of the present invention have been described in detail above, but the present invention is not limited to the described embodiments. For those skilled in the art, various changes, modifications, substitutions, and variations made to these embodiments without departing from the principles and spirit of the present invention still fall within the protection scope of the present invention.

Claims

1. A method for enhancing and segmenting the edges of stroke lesions based on grouped depthwise convolution and residual connections, characterized in that, The method includes the following steps: S1, Data Preprocessing: Acquire multimodal MRI data, filter irrelevant regions according to the voxel truncation method, standardize the filtered multimodal MRI data, and divide it into training set, validation set and test set; S2, Component Encoder: Construct a 3D U-Net encoder with a five-level downsampling structure. Each level contains two 3×3×3 convolutional layers. Each convolution is followed by instance normalization and LeakyReLU activation. Downsampling is achieved through 3×3×3 convolutions with a stride of 2. The number of feature channels are 32, 64, 128, 256, and 320 respectively. S3, Shallow Edge Enhancement: A shallow edge enhancement module is introduced in the first two high-resolution stages of the encoder. Multi-directional edge responses are extracted through grouped deep convolutions and residual fusion is performed with the original features to enhance the network's sensitivity to lesion boundaries. S4, Decoding and Restoration: The decoder, which has a structure symmetrical to the encoder, gradually restores the spatial resolution of the feature map through transposed convolution, and fuses the features of the corresponding layers of the encoder through skip connections to restore the labeled MRI image.

2. The method according to claim 1, characterized in that, The shallow edge enhancement mentioned in step S3 specifically includes: S31. Let the input feature map be X∈R C×D×H×W Where C is the number of channels, and D, H, and W are the depth, height, and width, respectively; S32 extracts spatial edge features through 3×3×3 grouped depthwise convolutions. The number of groups is equal to the number of input channels C, and the padding is 1. Each group of convolutions independently learns the spatial edge pattern of the corresponding channel. S33, batch normalization and ReLU activation are performed sequentially on the output of the grouped depthwise convolution to accelerate convergence and introduce nonlinearity; S34, through residual connection, adds the output of S33 to the original input feature map X element by element to perform feature fusion, so as to preserve the original feature information and alleviate gradient vanishing; S35 uses 1×1×1 convolution to achieve cross-channel feature fusion, integrating and combining the spatial edge features extracted independently from each group, and outputting an enhanced feature map X_out of the same size. The overall calculation method is as follows: , Among them, Conv 3×3×3 `group` represents a 3×3×3 depthwise convolution with C groups, and `BN` represents batch normalization. 1×1×1 This represents a 1×1×1 convolution.

3. The method according to claim 2, characterized in that, In the shallow edge enhancement module, the kernel size of the grouped depth convolution is 3×3×3 with padding of 1. The number of groups is equal to the number of input channels C, so that each channel performs spatial convolution independently. The number of parameters is 1 / C of that of a regular 3×3×3 convolution. The residual connection adds the output of the grouped depth convolution path after batch normalization and ReLU activation to the original input feature map X element by element, and then completes cross-channel feature fusion through 1×1×1 convolution. This allows the output feature map to retain the original feature information while incorporating multi-directional edge enhancement features.

4. The method according to claim 3, characterized in that, The shallow edge enhancement module only operates on the first two high-resolution stages of the encoder, namely the stages with 32 and 64 channels, and does not operate on the deep low-resolution stages; the output of the shallow edge enhancement module is an enhanced feature map of the same size, which is then passed down to the next encoder level.

5. The method according to claim 4, characterized in that, In the grouped depthwise convolution, the convolution kernels of different channels spontaneously converge toward different spatial gradient directions during training, including horizontal, vertical and diagonal directions, thereby achieving parallel detection of multi-directional edges; the direction pattern is driven by data-driven adaptive learning rather than manually preset fixed directions.

6. The method according to claim 5, characterized in that, The input features received by the decoder include: the decoder's own features after upsampling, and the skip connection features from the corresponding level of the encoder. The first two skip connection features are features enhanced by the shallow edge enhancement module, which enables the decoder to use the enhanced edge information for more accurate boundary localization when restoring spatial resolution. The decoder output is mapped to the number of target categories by a 1×1×1 convolution, and a probability map of each voxel belonging to the lesion is generated by the Sigmoid activation function. The final segmentation result is obtained by taking a threshold of 0.

5.

7. The method according to claim 6, characterized in that, The encoder and decoder infrastructure follows the nnU-Net adaptive configuration framework, including automatic data resampling, intensity normalization, and patch size and batch size configuration based on GPU memory limitations.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.