TransNeXt-based wavelet transform stroke image segmentation method
By adopting TransNeXt aggregate attention blocks and two-dimensional wavelet transformation in the Swin-Unet model, the shortcomings of the Swin-Unet model in capturing high-frequency information are solved, and the accuracy and stability of stroke image segmentation are improved.
Patent Information
- Application Number
- CN202510137289.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-27
AI Technical Summary
When the Swin-Unet model processes brain tissue boundaries, it fails to effectively capture high-frequency information, resulting in undersegmentation and affecting the segmentation results.
The wavelet transform stroke image segmentation method based on TransNeXt is used to replace the Swin-Transformer block by TransNeXt aggregate attention block, and high-frequency and low-frequency information are extracted in combination with two-dimensional wavelet transform, and texture and contour information in the image are retained.
It improves the accuracy of image segmentation, avoids model overfitting, enhances performance at the edges, and can more accurately segment the boundaries of the lesions.
Smart Images

Figure CN120047456A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image segmentation, and particularly to a method for segmenting stroke images based on wavelet transform of TransNeXt. Background Art
[0002] Ischemic Stroke is a disease in clinical medicine, and its cause is often the reduction of blood flow in the brain leading to hypoxia and damage of brain tissue. In this context, the segmentation of ischemic stroke lesions has become an important downstream task in medical imaging, and this task aims to accurately locate the area of infarction lesions in the patient's brain.
[0003] The traditional method in medical imaging is visual inspection. However, this method is affected by individual differences and experience levels of doctors. Therefore, the detection results generated by using this method have low accuracy, slow speed and poor stability, which affect the efficiency and accuracy of clinical decision-making. In recent years, with the rapid development of deep learning, the image segmentation model based on the U-Net architecture has received extensive attention due to its excellent performance in medical image segmentation.
[0004] The U-Net model can effectively extract multi-scale features of images and combine low-level features with high-level features through skip connections, so as to achieve accurate segmentation results.
[0005] The Swin-Unet model is a segmentation model based on Swin Transformer blocks and the U-Net architecture, which combines the powerful features of the U-Net with the sliding window self-attention mechanism of the Swin Transformer, aiming to improve the accuracy and robustness of medical image segmentation.
[0006] However, under-segmentation often occurs in Swin-Unet. The main reason for the under-segmentation in Swin-Unet is the insufficient accuracy of feature extraction in the Swin-Transformer blocks. Although Swin-Unet adopts a hierarchical design and gradually extracts features through deep learning, when dealing with the boundaries of brain tissues, the Swin-Transformer blocks fail to effectively capture the high-frequency information at the edges. These high-frequency information usually contains subtle structural differences, and these differences often can reflect the actual morphology and location of brain lesions.
[0007] In addition, in the subsequent downsampling step, the Swin-Unet model performs Patch Merge and reduces the resolution of the feature map. At this time, if high-frequency information that was not captured during the upstream feature extraction process, the high-frequency information may be permanently lost due to downsampling. The loss of high-frequency information causes the model to perform poorly at the edges, unable to accurately segment the boundaries of the lesions, resulting in under-segmentation and ultimately affecting the segmentation result. Summary of the Invention
[0008] To solve the technical problem of poor accuracy in segmenting ischemic stroke lesion images, the present invention proposes a method for segmenting stroke images based on TransNeXt wavelet transform.
[0009] The present invention provides a method for segmenting stroke images based on TransNeXt wavelet transform, which includes:
[0010] Obtain the image to be segmented, and input the image to be segmented into a pre-trained image segmentation model. Through the image segmentation model, obtain the segmentation result, where the image segmentation model includes: a preprocessing module, an encoder module, a low-frequency feature extraction module, a high-frequency feature extraction module, and a decoder module;
[0011] The training process of the image segmentation model includes:
[0012] Through the preprocessing module, preprocess the ischemic stroke lesion image;
[0013] Input the preprocessed ischemic stroke lesion image into the high-frequency feature extraction module for high-frequency feature information extraction;
[0014] Input the preprocessed ischemic stroke lesion image into the encoder module for feature extraction;
[0015] Input the output result of the encoder module into the low-frequency feature extraction module for low-frequency feature information extraction;
[0016] The decoder module reconstructs the lesion image based on the features extracted by the encoder module and the low-frequency feature extraction module, and obtains the segmentation result based on the high-frequency feature extraction module to complete the training of the image segmentation model.
[0017] Optionally, the step of inputting the preprocessed ischemic stroke lesion image into the high-frequency feature extraction module for high-frequency feature information extraction includes:
[0018] Use the pywt.dwt2 operation to obtain the high-frequency components Fi H , Fi V , Fi D ;
[0019] The three obtained high-frequency components are merged, and after passing through a linear layer and normalization, a high-frequency feature map Fi is obtained. HF 。
[0020] Optionally, inputting the preprocessed ischemic stroke lesion image into the encoder module for feature extraction includes:
[0021] The encoder module consists of a Patch Embed, a stage1 module, a stage2 module, and a stage3 module; the stage1 module, stage2 module, and stage3 module of the encoder module are all composed of a first TransNeXt block, a second TransNeXt block, and a patch merging layer;
[0022] Input the preprocessed ischemic stroke lesion image Fi into the Patch Embed to obtain non-overlapping image patches of 4×4;
[0023] Input the non-overlapping image patches of 4×4 into the patch embedding layer to obtain the feature map Fi 1 ;
[0024] Input the feature map Fi 1 into the stage1 module of the encoder module, and output to obtain the feature map Fi 2 ;
[0025] Input the feature map Fi 2 into the stage2 module of the encoder module, and output to obtain the feature map Fi 3 ;
[0026] Input the feature map Fi 3 into the stage3 module of the encoder module, and output to obtain the feature map Fi 4 。
[0027] Optionally, preprocessing the ischemic stroke lesion image through the preprocessing module includes:
[0028] Cut the original file to obtain an image group and a corresponding label group and perform screening;
[0029] Perform image enhancement on the screened image data;
[0030] Cut the image from the center of the picture, and the size of the cut picture is 224×224;
[0031] Perform attention masking on the picture data, with the masking rate set to 0.1 and the masking block size set to 7×7;
[0032] Divide the picture data into a training set and a test set, with a ratio of 4:1;
[0033] Convert the image data into a tensor format, normalize it, and limit the range of tensor element values to [0, 1] to obtain the preprocessed ischemic stroke lesion image Fi.
[0034] Optionally, Patch Embed in the encoder module implements the following steps:
[0035] Use a convolutional kernel with a size of 4×4 and a stride of 4 to perform Patch Partition and Position Embed operations on the preprocessed ischemic stroke lesion image Fi to reduce the size of the feature map to the original The size is [1, 72, 56, 56];
[0036] Input the feature map into the LN layer for Layer Norm operation;
[0037] Perform Drop out operation on the feature map to obtain the feature map Fi 1 ;
[0038] Fi 1 Generate query Q, key K, and value V through a linear layer, and perform L2 normalization and scaling on the query; calculate the local similarity Attn pool and the pooled feature Fi 1 pool ; Calculate the pooled similarity Attn pool Combine the local and pooled similarities to calculate the attention weight Attn; obtain the final output by weighted aggregation of the local and pooled features, and generate the feature representation Fi through the Linear layer and the Dropout layer aggre ; Pass through the Convolution GLU layer to obtain the feature representation
[0039] Optionally, inputting the output result of the encoder module into the low-frequency feature extraction module for low-frequency feature information extraction includes:
[0040] The low-frequency feature extraction module consists of the pywt.dwt2 operation, Convolution GLU, and a linear layer;
[0041] Input the feature map Fi 2 、Fi 3 、Fi 4 into the low-frequency feature extraction module, and output to obtain the low-frequency feature map
[0042] Optionally, the decoder module reconstructs the lesion image based on the features extracted by the encoder module and the low-frequency feature extraction module, and obtains the segmentation result based on the high-frequency feature extraction module, including:
[0043] The decoder module consists of a Bottleneck module, a stage1_up module, a stage1_up module, a stage2_up module, a stage3_up module, and a low-frequency feature extraction module; the Bottleneck module consists of a first TransNeXt module and a second TransNeXt module; the stage1_up module, the stage2_up module, and the stage3_up module are all composed of a patch expansion layer, a first TransNeXt module, and a second TransNeXt module;
[0044] Input the feature map Fi 4 into the Bottleneck module, and output the feature map Fi 5 ;
[0045] Input the feature map Fi 5 and merge them, and input the result into the stage1_up module, and output the feature map Fi 6 ;
[0046] Input the feature map Fi 6 and merge them, and input the result into the stage2_up module, and output the feature map Fi 7 ;
[0047] Input the feature map Fi 7 and merge them, and input the result into the stage3_up module, and output the result through the Patch_Expanding layer to obtain the feature map Fi 8 ;
[0048] Input the feature map Fi 8 and Fi HF merge them, and after processing by Convolution GLU, obtain Fi res .
[0049] Optionally, train and test based on the training set and the test set, and evaluate the segmentation result, including:
[0050] The formula corresponding to the loss function is: loss = α × DiceLoss + β × BCE WithLogits Loss, where α = 0.6 and β = 0.4;
[0051] Use the Loss to train the image segmentation model to obtain an optimized image segmentation model;
[0052] The Dice coefficient is an index used to evaluate the similarity between the predicted value and the true value, and its corresponding formula is:
[0053]
[0054] Among them, pred represents the prediction result, and true represents the correct result.
[0055] The present invention has the following beneficial effects:
[0056] The wavelet transform stroke image segmentation method based on TransNeXt of the present invention uses a TransNeXt aggregation attention block to replace the Swin-Transformer block to achieve high-precision feature extraction. Two-dimensional wavelet transform is used to extract the high-frequency and low-frequency information of the image, and the texture and contour information in the image are retained. Self-attention masks are used to preprocess the image to avoid the model from overpaying attention to the high-frequency information details in the picture, thereby avoiding model overfitting, and thus improving the accuracy of image segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0058] Figure 1 It is a schematic diagram of the usage process of the image segmentation model of the present invention;
[0059] Figure 2 It is a flowchart of the wavelet transform stroke image segmentation method based on TransNeXt of the present invention;
[0060] Figure 3 It is a schematic diagram of the structure of the high-frequency information extraction module of the present invention;
[0061] Figure 4 It is a schematic diagram of the structure of the encoder module of the present invention;
[0062] Figure 5 It is a schematic diagram of the structure of the TransNeXt module of the present invention;
[0063] Figure 6 It is a schematic diagram of the structure of the Aggregate Attention module of the present invention;
[0064] Figure 7 It is a schematic diagram of the structure of the low-frequency feature extraction module of the present invention;
[0065] Figure 8 It is a schematic diagram of the structure of the Convolution GLU layer of the present invention. Detailed implementation manners
[0066] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following will, in conjunction with the accompanying drawings and preferred embodiments, detail the specific implementation manners, structures, features and their effects of the technical solutions proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0067] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0068] The TransNeXt model generates local attention by introducing a sliding window mechanism, and then combines the global attention obtained through activation and pooling methods to obtain an attention mechanism that simulates biological vision, aiming to improve the information aggregation ability of the model.
[0069] In addition to deep learning models, the two-dimensional discrete wavelet transform method (2D Discrete Wavelet Transform) has the characteristics of multi-resolution analysis. It can decompose an image into sub-bands of different scales and directions, so as to more effectively extract and represent image features. When dealing with complex images containing noise and textures, the application of the two-dimensional discrete wavelet transform method effectively improves the reliability of the segmentation results.
[0070] The present invention provides a wavelet transform stroke image segmentation method based on TransNeXt, and this method includes the following steps:
[0071] Obtain the image to be segmented, and input the image to be segmented into a pre-trained image segmentation model. Through the image segmentation model, obtain the segmentation result. Among them, the image to be segmented can be an ischemic stroke lesion image to be segmented. The image segmentation model includes: a preprocessing module, an encoder module, a low-frequency feature extraction module, a high-frequency feature extraction module and a decoder module. The usage process of the image segmentation model is as Figure 1 shown.
[0072] Refer to Figure 2 , which shows the process corresponding to the training process of the image segmentation model. The training process of this image segmentation model includes the following steps:
[0073] Step S1, preprocess the ischemic stroke lesion image through the preprocessing module.
[0074] In some embodiments, data files in the nzz.gz format can be obtained from the public dataset ISLES2022. In the training set, n takes the value of 400, where n here refers to the number of samplings, that is, the number of samples in the training set can be 400; in the test set, n takes the value of 100, that is, the number of samples in the test set can be 100; using the os.listdir() function, tqdm.tqdm() function, niblabel.load() function, and numpy.uint8() function, image slice examples Fi are obtained, and the number of slices of a single nii.gz format file is 20.
[0075] As an example, this step may include the following steps:
[0076] First step, cut the original file to obtain an image group and a corresponding label group and perform screening.
[0077] Second step, perform image enhancement on the screened image data.
[0078] Third step, cut the image from the center of the picture, and the size of the cut picture is 224×224.
[0079] Fourth step, perform attention masking on the picture data, with the masking rate set to 0.1 and the masking block size set to 7×7.
[0080] Fifth step, divide the picture data into a training set and a test set, and the ratio is 4:1.
[0081] Sixth step, convert the picture data into the Tensor format and perform normalization, restricting the range of tensor element values to [0, 1] to obtain the preprocessed ischemic stroke lesion image Fi.
[0082] Step S2, input the preprocessed ischemic stroke lesion image into the high-frequency feature extraction module for high-frequency feature information extraction.
[0083] In some embodiments, the preprocessed ischemic stroke lesion image can be input into the high-frequency feature extraction module for high-frequency feature information extraction. The structure of the high-frequency information extraction module is as Figure 3 shown.
[0084] As an example, this step may include the following steps:
[0085] First step, use the pywt.dwt2 operation to obtain the high-frequency components Fi H 、Fi V 、Fi D .
[0086] Optionally, the wavedec2 method can be used for Fi to obtain the high-frequency components Fi in the horizontal, vertical, and diagonal directions respectively H 、Fi V 、Fi D 。
[0087] In the second step, the three obtained high-frequency components are merged, and after passing through a linear layer and normalization, a high-frequency feature map Fi is obtained HF ,and its corresponding formula can be: Fi HF =Norm(Linear(cat(dwt2(Fi))))。
[0088] Optionally, the torch.cat operation can be performed on the above high-frequency components, and after passing through a linear layer and an interpolation operation, a feature map Fi is obtained HF ,with a size of [1, 3, 224, 224].
[0089] Step S3, input the preprocessed ischemic stroke lesion image into the encoder module for feature extraction.
[0090] Among them, the encoder module (Encoder) consists of Patch Embed, stage1 module, stage2 module, and stage3 module. The stage1 module, stage2 module, and stage3 module of the encoder module are all composed of the first TransNeXt block, the second TransNeXt block, and the Patch Merge layer. That is, each stage module is composed of 2 layers of TransNeXt modules and Patch Merge. The structure of the encoder module can be as Figure 4 shown. The structure of the TransNeXt module can be as Figure 5 shown.
[0091] As an example, inputting the preprocessed ischemic stroke lesion image into the encoder module for feature extraction can include the following steps
[0092] In the first step, input the preprocessed ischemic stroke lesion image Fi into Patch Embed to obtain non-overlapping image patches of 4×4.
[0093] In the second step, input the non-overlapping image patches of 4×4 into the patch embedding layer to obtain a feature map Fi 1 ,and its corresponding formula can be: Fi 1 =transpose(flatten(Conv(Fi))).
[0094] In the third step, input the feature map Fi 1 into the stage1 module of the encoder module, and output to obtain a feature map Fi2 , and its corresponding formula can be: Fi 2 = Patch_Merging(TransNeXt(TransNeXt(Fi))).
[0095] Fourth, input the feature map Fi 2 into the stage2 module of the encoder module, and the output is the feature map Fi 3 , and its corresponding formula can be: Fi 3 = Patch_Merging(TransNeXt(TransNeXt(Fi 2 ))).
[0096] Fifth, input the feature map Fi 3 into the stage3 module of the encoder module, and the output is the feature map Fi 4 , and its corresponding formula can be: Fi 4 = Patch_Merging(TransNeXt(TransNeXt(Fi 3 ))).
[0097] Optionally, Patch Embed in the encoder module (Encoder) can implement the following steps:
[0098] First, use a convolutional kernel with a size of 4×4 and a stride of 4 to perform Patch Partition and Position Embed operations on the preprocessed ischemic stroke lesion image Fi, so that the size of the feature map is reduced to a size of [1, 72, 56, 56].
[0099] Second, input the feature map into the LN layer for Layer Norm operation.
[0100] Third, perform a Drop out operation on the feature map to obtain the feature map Fi 1 .
[0101] Fourth, the structure of the Aggregate Attention module in stage1 is as Figure 6 shown: Fi 1 First, generate query Q, key K, and value V through a linear layer, and perform L2 normalization and scaling on the query. Subsequently, calculate the local similarity Attn pool and the pooled feature Fi 1 pool . Then, calculate the pooled similarity and calculate the attention weight Attn by combining the local and pooled similarities 1。Then, by aggregating local and pooling features with weights, the final output is obtained, and through the Linear layer and Dropout layer, a feature representation is generated After passing through the Convolution GLU layer, a feature representation is obtained
[0102] For example, the steps performed by Patch Embed are as follows:
[0103]
[0104] K, V = split(kv(Fi 1 ));
[0105] Attn local = (Q·K T + B local )⊙M; B local is the local bias, and M is the mask;
[0106] B pool is the global pooling bias;
[0107] Fi 1 local = (B local + Attn local )·V local ;
[0108] Fi 1 pool = A pool ·V pool ;
[0109]
[0110] For the above feature map Fi 1 GLU Perform Patch Merge to obtain the feature map Fi 2 , with a size of [1, 144, 28, 28];
[0111] In stage2, repeat the steps performed by Patch Embed to obtain the feature map Fi 3 , with a size of [1, 288, 14, 14];
[0112] In stage3, repeat the steps performed by Patch Embed to obtain the feature map Fi 4 , with a size of [1, 576, 7, 7].
[0113] Step S4: Input the output result of the encoder module into the low-frequency feature extraction module to extract low-frequency feature information.
[0114] In some embodiments, a low-frequency feature extraction module is constructed between the Encoder and the Decoder module to extract the low-frequency image information retained in the feature map, and the feature maps of each layer are used for low-frequency feature extraction. The structure of the low-frequency feature extraction module is as Figure 7 shown.
[0115] Among them, the low-frequency feature extraction module consists of the pywt.dwt2 operation, Convolution GLU, and linear layers.
[0116] As an example, this step may include the following steps:
[0117] First step, retain the output feature maps Fi 2 , Fi 3 , Fi 4 of stage1, stage2, and stage3. Using haar as the wavelet basis and the decomposition level of 2, use wavedec2(wavelet='haar', level = 2) to extract low-frequency features in the channel dimension That is, the feature maps Fi 2 , Fi 3 , Fi 4 can be input into the low-frequency feature extraction module, and the output is the low-frequency feature map The corresponding formulas are respectively:
[0118]
[0119] Second step, the structure of the Convolution GLU layer is as Figure 8 shown. Through the Convolution GLU layer and the linear layer, extract low-frequency features and use interpolation to convert the size of the feature map back to the input size to obtain the low-frequency feature map
[0120] Step S5, the decoder module reconstructs the lesion image based on the features extracted by the encoder module and the low-frequency feature extraction module, and obtains the segmentation result based on the high-frequency feature extraction module to realize the training of the image segmentation model.
[0121] Among them, the decoder module (Decoder) consists of a Bottleneck module, a stage1_up module, a stage1_up module, a stage2_up module, a stage3_up module, and a low-frequency feature extraction module; the Bottleneck module consists of a first TransNeXt module and a second TransNeXt module; the stage1_up module, the stage2_up module, and the stage3_up module are each composed of a patch expansion layer, a first TransNeXt module, and a second TransNeXt module. That is, each stage_up module consists of 2 layers of TransNeXt modules and Patch Expand.
[0122] As an example, the decoder module reconstructs the lesion image based on the features extracted by the encoder module and the low-frequency feature extraction module, and the segmentation result obtained based on the high-frequency feature extraction module may include the following steps:
[0123] First step, input the feature map Fi 4 output by the stage3 module into the Bottleneck module, and the output is the feature map Fi 5 , and its corresponding formula can be: Fi 5 = TransNext(TransNext(Fi 4 ))
[0124] It should be noted that in the stage1_up module: similar to the stage1 module, the feature map Fi 5 and the low-frequency feature map undergo a torch.cat operation, and query Q, key K, and value V are generated through a linear layer, and the query is L2-normalized and scaled. Subsequently, the local similarity and the pooled feature are calculated. Then, the pooled similarity is calculated, and the attention weight Attn 5 is calculated by combining the local and pooled similarities. Then, the local and pooled features are weighted and aggregated to obtain the final output, and through the Linear layer and the Dropout layer, a feature representation is generated. Finally, through the Convolution GLU layer, a feature representation
[0125] Second step, merge the feature map Fi 5 with , and input it into the stage1_up module, and the output result is the feature map Fi 6 , and its corresponding formula can be:
[0126]
[0127] In the third step, combine the feature map Fi 6 with and input it into the stage2_up module. The output result is the feature map Fi 7 , and its corresponding formula can be:
[0128]
[0129] In the fourth step, combine the feature map Fi 7 with and input it into the stage3_up module. Then, output the result through the Patch_Expanding layer to obtain the feature map Fi 8 , and its corresponding formula can be:
[0130]
[0131] In the fifth step, after combining the feature map Fi 8 with Fi HF , perform Convolution GLU processing to obtain Fi res , and its corresponding formula can be: Fi res = GLU(concat(Fi 8 , Fi HF ))). That is, after performing the torch.cat operation on Fi HF and Fi 8 , process the obtained feature map through Convolution GLU to obtain the final segmentation result Fi res , with the size of [1, 3, 224, 224].
[0132] Optionally, based on the training set and the test set, perform training and testing. Evaluating the segmentation result may include the following steps:
[0133] In the first step, use the Adam optimizer for training. Among them, the learning rate is set to 0.0001, and other parameters follow the default settings of Swin-Unet. A total of 300 epochs are trained.
[0134] In the second step, the loss function combines BCE With Logits Loss and Dice Loss, and the weights are set to 0.6 and 0.4 respectively. The formula corresponding to the loss function is: loss = α × DiceLoss + β × BCE WithLogits Loss, where α = 0.6 and β = 0.4.
[0135] In the third step, use the Loss to train the image segmentation model to obtain an optimized image segmentation model.
[0136] In the fourth step, the Dice coefficient is an index used to evaluate the similarity between the predicted value and the true value, and its corresponding formula can be:
[0137]
[0138] where pred represents the prediction result and true represents the correct result.
[0139] In summary, the present invention uses the TransNeXt aggregation attention block to replace the Swin-Transformer block to achieve high-precision feature extraction. The two-dimensional wavelet transform is used to extract the high-frequency and low-frequency information of the image and retain the texture and contour information in the image. The self-attention mask is used to preprocess the image to avoid the model overpaying attention to the details of the high-frequency information in the picture, thereby avoiding model overfitting and improving the accuracy of image segmentation.
[0140] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention and should all be included within the protection scope of the present invention.
Claims
1. A wavelet transform stroke image segmentation method based on TransNeXt, characterized in that: The following steps are involved: Obtain an image to be segmented, and input the image to be segmented into a pre-trained image segmentation model, and obtain a segmentation result through the image segmentation model, wherein the image segmentation model includes: a preprocessing module, an encoder module, a low-frequency feature extraction module, a high-frequency feature extraction module and a decoder module; The training process of the image segmentation model includes: The ischemic stroke lesion images are preprocessed through the preprocessing module; The preprocessed ischemic stroke lesion image is input into the high-frequency feature extraction module to extract high-frequency feature information; The preprocessed ischemic stroke lesion image is input into the encoder module for feature extraction; The output result of the encoder module is input into the low-frequency feature extraction module to extract low-frequency feature information; The decoder module reconstructs the lesion image based on the features extracted by the encoder module and the low-frequency feature extraction module, and obtains the segmentation result based on the high-frequency feature extraction module to realize the training of the image segmentation model.
2. The method for stroke image segmentation based on wavelet transform of TransNeXt according to claim 1, characterized in that: The step of inputting the preprocessed ischemic stroke lesion image into a high-frequency feature extraction module to extract high-frequency feature information includes: Use pywt.dwt2 to obtain the high-frequency component Fi of the preprocessed ischemic stroke lesion image Fi H Fi V Fi D ; The three high-frequency components are combined and processed by linear layer and normalization to obtain the high-frequency feature map Fi HF .
3. The method for stroke image segmentation based on wavelet transform of TransNeXt according to claim 1, characterized in that: The step of inputting the preprocessed ischemic stroke lesion image into an encoder module for feature extraction comprises: The encoder module consists of Patch Embed, stage1 module, stage2 module and stage3 module; the stage1 module, stage2 module and stage3 module of the encoder module are all composed of the first TransNeXt block, the second TransNeXt block and the patch merging layer; The preprocessed ischemic stroke lesion image Fi is input into Patch Embed to obtain a 4×4 non-overlapping image block; Input the 4×4 non-overlapping image blocks into the patch embedding layer to obtain the feature map Fi 1 ; The feature map Fi 1 Input to the stage1 module of the encoder module, and output the feature map Fi 2 ; The feature map Fi 2 Input to the stage2 module of the encoder module, and output the feature map Fi 3 ; The feature map Fi 3 Input to the stage3 module of the encoder module, and output the feature map Fi 4 .
4. The method for stroke image segmentation based on wavelet transform of TransNeXt according to claim 1, characterized in that: The preprocessing module is used to preprocess the ischemic stroke lesion image, including: Cut the original file to obtain image groups and corresponding label groups and filter them; Performing image enhancement on the filtered image data; The image is cut from the center of the image, and the size of the cut image is 224×224; Perform attention masking on the image data, set the masking rate to 0.1, and set the masking block size to 7×7; Divide the image data into training set and test set with a ratio of 4:1; The image data is converted into a tensor format and normalized, and the value range of the tensor elements is limited to [0, 1] to obtain the preprocessed ischemic stroke lesion image Fi.
5. The method for stroke image segmentation based on wavelet transform of TransNeXt according to claim 3 is characterized in that: Patch Embed in the encoder module implements the following steps: The convolution kernel with a size of 4×4 and a step size of 4 is used to perform PatchPartition and Position Embed operations on the preprocessed ischemic stroke lesion image Fi to reduce the size of the feature map to the original size. The size is [1, 72, 56, 56]; Input the feature map to the LN layer for Layer Norm operation; The feature map is dropped out to obtain the feature map Fi 1 ; Fi 1 Generate query Q, key K and value V through linear layer, and perform L2 normalization and scaling on the query; Calculate the local similarity Attn pool And pooling feature Fi 1 pool ; Calculate the pooled similarity Attn pool The attention weight Attn is calculated by combining the local and pooled similarities; the final output is obtained by weighted aggregation of local and pooled features, and the feature representation Fi is generated through the Linear layer and the Dropout layer. aggre ; After the Convolution GLU layer, the feature representation is obtained 6. The method for stroke image segmentation based on wavelet transform of TransNeXt according to claim 3, characterized in that: The step of inputting the output result of the encoder module into the low-frequency feature extraction module to extract low-frequency feature information includes: The low-frequency feature extraction module consists of pywt.dwt2 operations, Convolution GLU and linear layers; The feature map Fi 2 Fi 3 Fi 4 Input to the low-frequency feature extraction module and output the low-frequency feature map 7. The method for stroke image segmentation based on wavelet transform of TransNeXt according to claim 6, characterized in that: The decoder module reconstructs the lesion image based on the features extracted by the encoder module and the low-frequency feature extraction module, and obtains the segmentation result based on the high-frequency feature extraction module, including: The decoder module is composed of a Bottleneck module, a stage1_up module, a stage2_up module, a stage3_up module, and a low-frequency feature extraction module; the Bottleneck module is composed of a first TransNeXt module and a second TransNeXt module; the stage1_up module, the stage2_up module, and the stage3_up module are all composed of a patch expansion layer, a first TransNeXt module, and a second TransNeXt module; The feature map Fi 4 Input to the Bottleneck module and output the feature map Fi 5 ; The feature map Fi 5 and Merge and input to the stage1_up module, and output the result to get the feature map Fi 6 ; The feature map Fi 6 and Merge and input to the stage2_up module, and output the result to get the feature map Fi 7 ; The feature map Fi 7 and Merge and input to the stage3_up module, and output the result through the Patch_Expanding layer to obtain the feature map Fi 8 ; The feature map Fi 8 With Fi HF After merging, it is processed by Convolution GLU to obtain Fi res .
8. The method for stroke image segmentation based on wavelet transform of TransNeXt according to claim 4, characterized in that: Perform training and testing based on the training set and test set, and evaluate the segmentation results, including: The formula corresponding to the loss function is: loss = α × DiceLoss + β × BCE WithLogits Loss, where α = 0.6, β = 0.4; Use Loss to train the image segmentation model and obtain the optimized image segmentation model; The Dice coefficient is an indicator used to evaluate the similarity between the predicted value and the true value. The corresponding formula is: Among them, pred represents the predicted result and true represents the correct result.