A Stroke Lesion Segmentation Method Based on Adaptive Feature Enhancement
By designing adaptive feature enhancement technology in the stroke lesion segmentation method, and using iterative combined attention module and cross-level feature enhancement module for feature fusion, the problem of loss of feature information and low segmentation accuracy in the existing technology is solved, and higher segmentation accuracy and adaptability are achieved.
Patent Information
- Application Number
- CN202411113964.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-14
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-08-14
AI Technical Summary
The prior art has problems of loss of feature information and low segmentation accuracy in the image segmentation task of stroke lesion areas, especially when dealing with small target areas and heterogeneous lesions.
A stroke lesion segmentation method based on adaptive feature enhancement is designed, multi-level feature extraction is performed through an encoder, and an iterative combined attention module and cross-level feature enhancement and convergence module are used to perform adaptive and efficient fusion of features.
It effectively solves the problem of loss of feature information and low segmentation accuracy, improves the segmentation accuracy of small target areas and heterogeneous lesions, and enhances the adaptive ability of the model.
Smart Images

Figure CN119068191B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of stroke lesion segmentation based on deep learning, and in particular to a stroke lesion segmentation method based on adaptive feature enhancement. Background Art
[0002] As a disease that frequently occurs in clinical practice and has extremely serious consequences, stroke is characterized by a high disability rate and a high mortality rate. Clinically, diagnosing stroke usually requires doctors to observe the patient's brain CT (Computerized Tomography) or MRI (Magnetic Resonance Imaging). A group of brain MRIs often has a large number of slices, and it depends on the experience of professional doctors to judge whether there are corresponding lesions. In addition, some noises are generated during the imaging of medical images, resulting in mottled, granular, textured or snowflake-like appearances of the images. To address the above problems, the use of image automatic segmentation technology can assist in improving the diagnostic efficiency and accuracy of doctors.
[0003] In recent years, with the development of deep learning technology, architectures such as U-Net (U-shaped Network), DeepLab3+ (an encoder-decoder network for semantic image segmentation based on atrous separable convolution), and CLCI-Net (Cross-level fusion and Context Inference Networks, an image segmentation network for cross-level fusion and context inference) have shown obvious effects in improving the processing speed and accuracy of clinical medical images. However, in the image segmentation task of stroke lesion areas, there are still some deficiencies and limitations, bringing some problems and challenges. First, the lesions caused by stroke usually show a relatively limited distribution in the dataset. Although a group of brain MRI slices cover a large number of images, most of them are MRI images of healthy brains, resulting in a relatively small number of samples with lesion areas. Second, the pathological images of stroke patients show diverse and heterogeneous characteristics, that is, the sizes of the lesion areas vary greatly. For small target areas, due to multiple downsamplings, existing models cannot well obtain multi-scale feature information, resulting in a large loss of feature information, and thus cannot accurately segment small target areas, causing many traditional image segmentation methods to perform poorly in distinguishing lesion areas from normal tissues. Summary of the Invention
[0004] To address the above technical problems, the present invention proposes a stroke lesion segmentation method based on adaptive feature enhancement. This method extracts multi-level features through an encoder, and designs an iterative joint attention module and a cross-level feature reinforcement aggregation module for adaptive and efficient fusion of features; the extracted features are further decoded into the image space through a decoder to obtain the segmentation result of the lesion. The present invention includes two newly designed modules: an iterative attention residual block and a cross-level feature reinforcement aggregation module. For the iterative attention residual block, this module is the basic building unit of the encoder. Each iterative attention residual block contains 3 convolutional residual blocks, 1 iterative joint attention module, and 1 ReLu activation layer. Among them, the convolutional residual block can effectively extract multi-level features in the image, and the iterative joint attention module adaptively fuses features by iteratively combining global and local attention. Since a deeper encoder will reduce the resolution and cause loss of detail information, the iterative attention residual block uses the attention mechanism as a residual connection. Under the influence of the attention mechanism, more rich detail information is fused into the deep features, and more spatial information is incorporated into the deep features dominated by semantic information, thereby reducing the loss of detail information.
[0005] For the cross-level feature reinforcement aggregation module, this module uses information entropy to measure the amount of information of each channel feature, screens the shallow features of the skip connection, reduces the weight of a large amount of background information in the shallow features, reduces the impact of this information on the segmentation of the target area, constructs enhanced features, and solves the problem of low efficiency of feature reuse and fusion.
[0006] The present invention combines iterative joint attention and information entropy, so that the deep features contain rich and abstract semantic information while also containing relatively detailed spatial feature information. Using information entropy to measure the intensity of information in the features, the features are efficiently fused, thereby adaptively extracting important features in the image, and then effectively solving the problems existing in the stroke image segmentation task.
[0007] The specific steps of the method of the present invention include:
[0008] S1 Preprocess the MRI magnetic resonance imaging data: Scale it to a unified spatial resolution through linear interpolation as the image to be segmented.
[0009] S2 Input the image to be segmented into the encoder for multi-level feature extraction. The encoder includes a primary feature extraction module, three encoding modules, and two cross-level feature reinforcement aggregation modules CFA. The specific steps include:
[0010] S21 Input the image to be segmented into the primary feature extraction module for feature extraction. The primary feature extraction module is composed of a convolutional layer and a max pooling layer in cascade, and is used to extract the primary feature F of the image 0 .
[0011] S22 feeds the feature F extracted in step S21 0 , denoted as X, into the first encoding module. Each encoding module is composed of several cascaded iterative attention residual blocks. Each group of iterative attention residual blocks is composed of 3 convolutional residual blocks, 1 iterative joint attention module, and 1 ReLu activation layer cascaded. For the i-th encoding module, its input feature is F i-1 , and its output feature is F i .
[0012] The specific steps of the encoding module include:
[0013] S221 feeds the input feature X into the first convolutional residual block. The feature obtained by sequentially passing through the convolutional layer, batch normalization layer, and ReLu activation layer is added to the input feature residually to obtain the output feature of the first convolutional residual block
[0014] S222 feeds into the second convolutional residual block to obtain the output feature Feeds into the third convolutional residual block to obtain the output feature Among them, the calculation operations of the second and third convolutional residual blocks are the same as the calculation step S221 of the first convolutional residual block.
[0015] S223 feeds the input feature of this group of iterative attention residual blocks, that is, the input feature X of step S22, and the output feature of the third group of convolutional residual blocks , denoted as Y, into the iterative joint attention module to obtain the output feature Z. This module includes the following specific steps:
[0016] S2231 adds the input features X and Y by broadcasting to obtain the feature x, and feeds the feature x into the two branches of the first global-local joint attention module: the global attention branch and the local attention branch.
[0017] S2232 For the global attention branch, the input feature x first passes through the global average pooling layer f AP (), then enters the first pointwise convolutional layer f PW1 (), then passes through the batch normalization layer f BN (), and the ReLu activation layer f ReLu (), and then enters the second pointwise convolutional layer f PW2 (), and finally passes through the batch normalization layer f BN (), and the Sigmoid activation layer f Sigmoid () to calculate the global attention A global .
[0018] For the local channel attention branch package, the input feature x first enters the first pointwise convolution layer f PW1 (), and then passes through the batch normalization layer f BN (), and the ReLu activation layer f ReLu (), and after calculation, it enters the second pointwise convolution layer f PW2 (), and finally passes through the batch normalization layer f BN (), and the Sigmoid activation layer f Sigmoid (), and the local attention A is calculated local .
[0019] S2234 Multiply the global attention A global and the local attention A local element-wise with the input features X and Y respectively, and add the two product results to obtain the feature x'.
[0020] S2235 Input the output feature x' of step S2234 into the second global-local joint attention module, repeat step S2232 to obtain the global attention A global′ , repeat step S2233 to obtain the local attention A local′ , and perform weighted addition on the input features X and Y of step S223 respectively to obtain the output Z of the second global-local joint attention module
[0021] S224 Perform ReLu activation operation on the feature Z to obtain the output feature of this group of iterative attention residual blocks
[0022] S225 Repeat steps S221 to S224 several times, and the number of repetitions is the number of groups of iterative attention residual blocks in the encoding module. The output feature of the last group of iterative attention residual blocks is the output feature of the encoding module. For the i-th encoding module, the input feature of its first group of iterative attention residual blocks is F i-1 , and the input features of other groups of iterative attention residual blocks are the output features of the previous group of iterative attention residual blocks
[0023] S23 Input the output feature F 1 of the first encoding module into the second encoding module, and calculate the output feature F 2 of the second encoding module; input the output feature map F 2 of the second encoding module into the third encoding module, and calculate the output feature map F 3 of the third encoding module. Among them, the operations of the second and third encoding modules are the same as those of the first encoding module, that is, step S22
[0024] S24 Input the input feature F 0(Shallow features) and output feature F 1 (Deep features) are input into the first cross-level feature enhancement aggregation module (CFA) to obtain the enhanced aggregated features The specific steps of this module are as follows:
[0025] S241 Calculate the channel information entropy E of the shallow features i 。
[0026] S242 For the shallow features, each channel feature f i is multiplied by the channel information entropy E i to obtain the entropy-weighted features
[0027] S243 Add the entropy-weighted features to the deep features to obtain the enhanced aggregated features
[0028] S25 Input the output feature F 2 (Shallow features) of the second encoding module and the output feature F 3 (Deep features) of the third encoding module into the second cross-level feature enhancement aggregation module to obtain the enhanced aggregated features The calculation steps of the second cross-level feature enhancement aggregation module are the same as those of the first cross-level feature enhancement aggregation module, namely steps S241 to S243
[0029] S3 Input the features f 3 、 and extracted by the encoder into the decoder to output the segmentation result map. The decoder is composed of three decoding modules and a segmentation head in cascade. Among them, each decoding module is composed of a transposed convolution layer, a normalization layer, and a Relu activation function layer in cascade
[0030] Advantages of the present invention: To solve the problems of poor real-time performance and insufficient feature extraction in the stroke lesion segmentation model, the present invention proposes a stroke lesion segmentation method based on adaptive feature enhancement. First, the iterative attention residual block in the encoder is used to extract features from the input stroke image, so that the deep features contain rich abstract semantic information and also rich spatial detail feature information. After feature extraction, a cross-level feature reinforcement aggregation module is used for efficient adaptive feature fusion. This module measures the amount of information in the features by calculating the channel information entropy, assigns a higher weight to the channels with rich information, and a lower weight to the channels with less information. The present invention has the following benefits: 1. An iterative joint attention module is designed. The iterative joint attention module can achieve channel attention at multiple scales by adjusting the size of the spatial pool, so as to be applicable to the segmentation of lesion areas of different sizes; 2. A cross-level feature reinforcement aggregation module is designed. By fusing features weighted by information entropy, the feature channels more beneficial to the final segmentation result are selected, making the segmentation result more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 Architecture diagram of the stroke lesion segmentation method based on adaptive feature enhancement;
[0032] Figure 2 Structural diagram of the iterative joint attention module;
[0033] Figure 3 Structural diagram of the cross-level feature reinforcement aggregation module;
[0034] Figure 4 Effect diagram of the stroke lesion segmentation by the technology of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0035] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0036] As a preferred implementation form of the present invention, a stroke lesion segmentation method based on adaptive feature enhancement is provided, and its architecture is as Figure 1 shown, including the following steps:
[0037] S1 Scale the MRI magnetic resonance imaging data to a unified 224×224 spatial resolution by linear interpolation method as the image to be segmented.
[0038] S2 Input the image to be segmented into the encoder for multi-level feature extraction. The encoder includes a primary feature extraction module, three encoding modules and two cross-level feature reinforcement aggregation modules (CFA). The specific steps include:
[0039] S21 inputs the image to be segmented into the encoder for multi-level feature extraction, and extracts the primary feature F of the image 0 , and the primary feature extraction module is composed of a cascaded convolutional layer and a max pooling layer. Among them, the convolutional kernel size of the convolutional layer is 7×7, the number of convolutional kernels is 64, the stride is 2, the padding is 3, and the convolutional kernel size of the max pooling layer is 3×3, the number of convolutional kernels is 64, the stride is 2, and the padding is 1.
[0040] S22 For the primary feature F 0 , also denoted as X, further extracts multi-level features through three encoding modules. The first encoding module is composed of three groups of attention residual blocks in cascade, and the second and third encoding modules are composed of four groups and six groups of attention residual blocks in cascade respectively. Each group of iterative attention residual blocks (steps S221 to S224) is composed of 3 convolutional residual blocks, 1 iterative joint attention module (whose structure is as Figure 2 shown) and 1 ReLu activation layer in series. For the i-th encoding module, its input feature is F i-1 , and its output feature is F i .
[0041] The specific steps of the encoding module include:
[0042] S221 inputs the feature X into the first convolutional residual block, and the feature obtained by calculating through the 1×1 convolutional layer, batch normalization layer, and ReLu activation layer is added to the input feature by residual, and the output feature of the first convolutional residual block is obtained
[0043] S222 inputs into the second convolutional residual block to obtain the output feature inputs into the third convolutional residual block to obtain the output feature Among them, the calculation steps of the second and third convolutional residual blocks (where the convolutional kernel sizes of the convolutional layers are 3×3 and 1×1 respectively) are the same as the calculation step S221 of the first convolutional residual block.
[0044] S223 inputs the input feature of this group of iterative attention residual blocks, that is, the input feature X of step S22, and the output feature of the third group of convolutional residual blocks , also denoted as Y, into the iterative joint attention module, and this module includes the following specific steps:
[0045] S2231 inputs the input features X and Y, and after broadcasting and adding, the feature x is obtained, and the feature x is respectively input into the two branches of the first global-local joint attention module: the global attention branch and the local attention branch.
[0046] The S2232 global attention branch contains a global average pooling layer, two pointwise convolutional layers, and two activation function layers. First, the input feature x passes through the global average pooling layer f AP (), and then enters the first pointwise convolutional layer f PW1 (). After that, it passes through the batch normalization layer f BN () and the ReLu activation layer f ReLu () and then enters the second pointwise convolutional layer f PW2 (). Finally, after passing through the batch normalization layer f BN () and the Sigmoid activation layer f Sigmoid (), the global attention A global is calculated. The specific calculation formula is:
[0047] A global = f Sigmoid (f BN (f PW2 (f ReLu (f BN (f PW1 (f AP (x)))))))
[0048] The global average pooling layer converts the feature x into a feature with a size of C×1×1, where C is the number of channels. The convolutional kernel sizes of f PW1 () and f PW2 () are (C / r)×C×1×1 and C×(C / r)×1×1 respectively, and the dimension of A global is C×1×1.
[0049] The S2233 local channel attention branch contains two pointwise convolutional layers and two activation function operations. First, the input feature x enters the first pointwise convolutional layer f PW1 (). After passing through the batch normalization layer f BN () and the ReLu activation layer f ReLu (), it enters the second pointwise convolutional layer f PW2 (). Finally, after passing through the batch normalization layer f BN () and the Sigmoid activation layer f Sigmoid (), the local attention A local is calculated. The specific calculation formula is:
[0050] A local = f Sigmoid (f BN (f PW2 (f ReLu (f BN (f PW1 (x))))))
[0051] f PW1( ) and f PW2 The convolution kernel sizes of ( ) are (C / r)×C×1×1 and C×(C / r)×1×1 respectively. A local The dimension of is C×H×W.
[0052] S2234 multiplies the global attention A global and the local attention A local with the input features X and Y element-wise respectively, adds the two resulting product results, and obtains the feature x′. The calculation formula is:
[0053]
[0054] The above X and Y are obtained from step S223. In the formula, represents element-wise multiplication.
[0055] S2235 inputs the output feature x′ of step S2234 into the second global-local joint attention module, repeats step S2232 to obtain the global attention A global′ , repeats step S2233 to obtain the local attention A local′ , and performs weighted addition on the feature maps X and Y input in step S223 respectively to obtain the output Z of the second global-local joint attention module. The specific formula is:
[0056]
[0057] S224 performs a ReLu activation operation on the feature Z to obtain the output feature of this group of iterative attention residual blocks.
[0058] S225 repeats steps S221 to S224 several times. The number of repetitions is the number of groups of iterative attention residual blocks in the encoding module (the first, second, and third encoding modules contain three, four, and six groups of iterative attention residual blocks respectively). The output feature of the last group of iterative attention residual blocks is the output feature of the encoding module. For the i-th encoding module, the input feature of its first group of iterative attention residual blocks is F i-1 , and the input features of other groups of iterative attention residual blocks are the output features of the previous group of iterative attention residual blocks.
[0059] S23 inputs the output feature F 1 of the first encoding module into the second encoding module, and calculates to obtain the output feature F 2 of the second encoding module; inputs the output feature map F 2 of the second encoding module into the third encoding module, and calculates to obtain the output feature map F 3 of the third encoding module. Among them, the steps of the second and third encoding modules are the same as those of the first encoding module, that is, step S22.
[0060] S24 inputs the input feature F 0 (shallow feature) and the output feature F 1 (deep feature) of the first coding module into the first cross-level feature enhancement aggregation module (CFA, whose structure is as Figure 3 shown), and obtains the enhanced aggregated feature The specific steps of this module are as follows:
[0061] S241 calculates the channel information entropy of the shallow feature, and the specific steps include:
[0062] S2411 For each channel feature f i , a binary tuple (C n , A n ) is extracted through a 3×3 sliding window, where C n represents the feature value at the center point of the nth 3×3 sliding window, and A n represents the average value of the remaining feature values in this window.
[0063] S2412 Counts the number of occurrences N(C n , A n ) of the binary tuple (C n , A n ), and calculates the probability P n , A n ) of the binary tuple (C n ), and its calculation formula is:
[0064] P n = N(C n , A n ) / (H×W)
[0065] where H is the height of the feature and W is the width of the feature.
[0066] S2413 Calculates the information entropy E i of the channel feature f i , and its calculation formula is:
[0067]
[0068] S242 For the shallow feature, each channel feature f i is multiplied by the channel information entropy E i to obtain the entropy-weighted feature, so as to focus on the channels with larger information entropy and obtain a richer information representation.
[0069] S243 Adds the entropy-weighted feature to the deep feature to obtain the enhanced aggregated feature. The feature obtained in this way has the spatial information of the shallow feature and at the same time contains the semantic information of the deep feature.
[0070] S25 feeds the output feature F of the second encoding module 2 (shallow feature) and the output feature F of the third encoding module 3 (deep feature) into the second cross-level feature enhancement aggregation module to obtain the enhanced aggregated feature The calculation steps of the second cross-level feature enhancement aggregation module are the same as those of the first cross-level feature enhancement aggregation module, namely steps S241 to S243.
[0071] S3 decodes the feature extracted by the encoder into the image space to obtain the lesion segmentation result. Specifically, the feature F extracted by the encoder 3 、 and are input into the decoder. The decoder is composed of three decoding modules and a segmentation head in cascade. Among them, each decoding module is composed of a transposed convolution layer, a normalization layer, and a Relu activation function in cascade. The specific steps of the decoder include:
[0072] S31 feeds the feature F 3 into decoding module 3, and calculates the decoded feature f through the transposed convolution layer, the normalization layer, and the Relu activation function layer 3 .
[0073] S32 cascades the feature and the feature f 3 in the channel dimension and then inputs them into decoding module 2, and calculates the decoded feature f through the transposed convolution layer, the normalization layer, and the Relu activation function layer 2 .
[0074] S33 cascades the feature and the feature f 2 in the channel dimension and then inputs them into decoding module 1, and calculates the decoded feature f through the transposed convolution layer, the normalization layer, and the Relu activation function layer 1 .
[0075] S34 feeds the feature f 1 into the segmentation head, and obtains the segmented image through the convolution layer and the Sigmoid activation layer, and its effect is as Figure 4 shown.
[0076] Example:
[0077] The steps of this example are the same as those in the specific implementation manner, and will not be elaborated here. The following shows some implementation processes and implementation results.
[0078] The technology of the present invention is implemented on the dataset ISLES2022, which contains 250 case images, including a large number of images of non-lesion regions. In order to reduce the impact of excessive negative samples on model training, when splitting the images, the middle slices were selected, and finally 15,684 pairs of images were obtained. Among them, 12,547 images were used as the training set, and 3,137 images were used as the test set. The above image data was cropped to a unified size of 224×224. This model was trained using the Adam (Adaptive Moment Estimation) optimizer with a batch size of 8 and a weight decay of 10 -8 , and the learning rate was set to 0.0001. In order to verify the effectiveness of the technology of the present invention, Figure 4 shows the segmentation effect diagram of the technology of the present invention after training on the dataset ISLES2022, and compares it with the following methods respectively: U-Net (U-shaped Network): an image segmentation network based on convolutional neural network, CA-Net (Comprehensive Attention Network): an image segmentation network based on comprehensive attention mechanism, DeepLab3+ (DeepLab Version 3Plus): an encoder-decoder network for semantic image segmentation based on dilated separable convolution, TransUNet (Transformers+U-Net): a medical image segmentation network based on Transformer and convolutional neural network, CLCI-Net (Cross-level fusion and Context Inference Networks): an image segmentation network for cross-level fusion and context inference, VA-TransUNet (Visual Attention with Transformers+U-Net): an image segmentation network using multi-scale visual attention. As Figure 4 shown, compared with methods such as U-Net and CLCI-Net, the technology of the present invention can accurately segment the lesions, especially the lesion regions with smaller areas, which can also be segmented by the technology of the present invention. Compared with other methods, the segmented lesion regions have clearer contours and are more similar to the labels in shape.
[0079] To verify the effectiveness of the technology of the present invention, Table 1 uses indicators such as Dice (measuring the overlap degree of two sets), Recall (recall rate, measuring the proportion of the number of samples correctly predicted as positive examples by the model in all actual positive example samples), VOE (Volume Overlap Error, measuring the error of the volume overlap degree between the segmentation result and the true segmentation), RVD (Relative Volume Difference, measuring the volume difference between the segmentation result and the true segmentation), and IoU (Intersection over Union, measuring the overlap degree between the prediction result and the true annotation) to compare the technology of the invention and the existing technology. As shown in Table 1, the technology of the present invention has achieved good results in multiple indicators, thus proving the effectiveness of the technology of the present invention.
[0080] Table 1 Quantitative comparison results of the technology of the invention and the existing technology
[0081]
Claims
1. A stroke lesion segmentation method based on adaptive feature enhancement, characterized in that: The specific steps include: S1. Preprocess the MRI data: scale the data to a uniform spatial resolution using a linear interpolation method as the image to be segmented; S2. Input the image to be segmented into the encoder for multi-level feature extraction, the encoder includes a primary feature extraction module, three encoding modules and two cross-level feature enhancement and aggregation modules CFA; The specific implementation process of the encoder is as follows: S21. Input the image to be segmented into the primary feature extraction module for feature extraction. The primary feature extraction module is composed of a convolution layer and a maximum pooling layer cascaded to extract the primary feature F0 of the image; S22. Input feature F0, denoted as X, to the first encoding module; each encoding module is composed of a cascade of several groups of iterative attention residual blocks, each group of iterative attention residual blocks is composed of a cascade of three convolutional residual blocks, an iterative joint attention module and a ReLu activation layer; for the i-th encoding module, its input feature is F i-1 , whose output feature is F i ; S23. Input the output feature F1 of the first encoding module to the second encoding module, and calculate the output feature F2 of the second encoding module; input the output feature map F2 of the second encoding module to the third encoding module, and calculate the output feature map F3 of the third encoding module. The operations of the second and third encoding modules are the same as those of the first encoding module; S24. Input the input feature F0 and output feature F1 of the first encoding module to the first cross-level feature enhancement aggregation module CFA to obtain the enhanced aggregation feature S25. Input the output feature F2 of the second encoding module and the output feature F3 of the third encoding module into the second cross-level feature enhancement aggregation module to obtain the enhanced aggregation feature The calculation steps of the second cross-level feature reinforcement aggregation module are the same as those of the first cross-level feature reinforcement aggregation module; S3. Input the output feature map of the encoding module in the encoder and the output feature map of the two cross-level feature enhancement aggregation modules CFA into the decoder, and output the segmentation result map. The decoder is composed of three decoding modules and a segmentation head in cascade.
2. The method for stroke lesion segmentation based on adaptive feature enhancement according to claim 1, characterized in that: The specific implementation process of the encoding module is as follows: S221. Input the input feature X to the first convolution residual block, and add the residuals of the features calculated by the convolution layer, batch normalization layer, and ReLu activation layer to the input features to obtain the output feature of the first convolution residual block S222. Input to the second convolution residual block to get the output features Will Input to the third convolution residual block to get the output features The computational operations of the second and third convolutional residual blocks are the same as the first convolutional residual block; S223. Combine the input features of this group of iterative attention residual blocks, i.e., feature X, with the output features of the third group of convolutional residual blocks Also denoted as Y, it is input into the iterative joint attention module to obtain the output feature Z; S224. Perform ReLu activation on feature Z to obtain the output features of this group of iterative attention residual blocks; S225. Repeat steps S221 to S224, the number of repetitions is the number of groups of iterative attention residual blocks in the encoding module, and the output features of the last group of iterative attention residual blocks are the output features of the encoding module; for the i-th encoding module, the input features of its first group of iterative attention residual blocks are F i-1 , the input features of other groups of iterative attention residual blocks are the output features of the previous group of iterative attention residual blocks.
3. The method for stroke lesion segmentation based on adaptive feature enhancement according to claim 2, characterized in that: The implementation process of the iterative joint attention module is as follows: S2231. Add the input features X and Y to obtain feature x, and input feature x into two branches of the first global-local joint attention module: a global attention branch and a local attention branch; S2232. For the global attention branch, the input feature x first passes through the global average pooling layer f AP (), then enter the first point-by-point convolution layer f PW1 (), and then through the batch normalization layer f BN () and ReLu activation layer f ReLu () and then enter the second point-by-point convolution layer f PW2 (), and finally through the batch normalization layer f BN () and Sigmoid activation layer f Sigmoid () Calculate the global attention A global ; S2233. For the local channel attention branch, the input feature x first enters the first point-by-point convolution layer f PW1 (), and then through the batch normalization layer f BN () and ReLu activation layer f ReLu () After calculation, enter the second point-by-point convolution layer f PW2 (), and finally through the batch normalization layer f BN () and Sigmoid activation layer f Sigmoid () Calculate the local attention A local ; S2234. Global attention A global and local attention A local Multiply the input features X and Y element by element, add the two product results, and get the feature x′; S2235 inputs feature x′ into the second global-local joint attention module and repeats step S2232 to obtain global attention A global′ , repeat step S2233 to get local attention A local′ , and perform weighted addition on features X and Y respectively to obtain the output feature Z of the second global-local joint attention module.
4. The method for stroke lesion segmentation based on adaptive feature enhancement according to claim 3, characterized in that: The specific implementation process of the cross-level feature enhancement aggregation module CFA is as follows: S241. Calculate the channel information entropy E of feature F0 i ; S242. For feature F0, each channel feature f i and channel information entropy E i Multiply them together to get the information entropy weighted features; S243. Add the information entropy weighted features to feature F1 to obtain enhanced aggregated features.
5. The method for stroke lesion segmentation based on adaptive feature enhancement according to claim 4, characterized in that: The specific calculation process of the channel information entropy of the feature is as follows: S2411. For each channel feature f i , extracting bigrams (C n ,A n ), C n represents the eigenvalue of the center point of the nth 3×3 sliding window, A n represents the average value of the remaining eigenvalues in the window; S2412. Calculate the pair (C n ,A n ) n , and its calculation formula is: P n =N(C n ,A n ) / (H×W) Among them, N(C n ,A n ) is a binary group (C n ,A n ) appears, H is the height of the feature, and W is the width of the feature; S2413. Calculate channel feature f i Information entropy E i , and its calculation formula is:
6. The method for stroke lesion segmentation based on adaptive feature enhancement according to claim 5, characterized in that: The specific implementation process of the decoder is as follows: S31 inputs the feature F3 into the decoding module 3, and obtains the decoded feature f3 through the deconvolution layer, the normalization layer and the Relu activation function layer; S32.Characteristics After being cascaded with feature f3 in the channel dimension, it is input into decoding module 2, and the decoded feature f2 is calculated through the deconvolution layer, normalization layer and Relu activation function layer; S33.Characteristics After being cascaded with feature f2 in the channel dimension, it is input into the decoding module 1, and the decoded feature f1 is calculated through the deconvolution layer, normalization layer and Relu activation function layer; S34. Input feature f1 into the segmentation head, and obtain the segmented image through the convolution layer and the Sigmoid activation layer.