Medical image segmentation method and system based on knowledge migration and attention mechanism

By constructing a medical image segmentation method based on knowledge transfer and attention mechanism, using two subnets and multi-scale spatial pixel attention modules, the problems of long-distance dependence and boundary blur in medical image segmentation are solved, and higher segmentation accuracy and robustness are achieved.

CN120339192APending Publication Date: 2025-07-18HENAN MECHANICAL & ELECTRICAL ENG COLLEGE
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510340255.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing deep learning-based medical image segmentation methods have difficulties in learning clear long-distance dependencies, and there are problems of blurred or discontinuous boundaries when marking segmentation boundaries.

Method used

Using a medical image segmentation method based on knowledge transfer and attention mechanism, by constructing two subnets with the same structure, the characteristics of the first subnet are migrated to the second subnet by using the knowledge transfer layer, and fused with their features, combining multi-scale spatial pixel attention modules and loss functions to improve the applicability and robustness of the network.

Benefits of technology

It effectively solves the problem of blurred boundaries of segmentation results, improves the accuracy and robustness of medical image segmentation, can better obtain long-distance dependencies and fine-grained local features, and generates clearer segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339192A_ABST
    Figure CN120339192A_ABST
Patent Text Reader

Abstract

The invention provides a medical image segmentation method and system based on knowledge migration and an attention mechanism, and belongs to the technical field of computer vision. The method comprises steps that a to-be-segmented medical image is segmented based on a trained medical image segmentation model, a segmentation result is acquired, and the medical image segmentation model comprises a first sub-network and a second sub-network; the first sub-network and the second sub-network adopt a network structure of an encoder and a decoder which are the same in structure; each first attention module of the encoder is in jump connection with a second attention module at a corresponding position of the decoder; a knowledge migration layer is arranged between the first attention module of the first sub-network and the first attention module at the same position of the second sub-network, and is used for taking first sub-features output by each first attention module of the first sub-network as knowledge to migrate to the second sub-network and performing feature fusion with the second sub-features; the method can effectively solve the problem that the boundary of the medical image segmentation result is fuzzy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and particularly to a medical image segmentation method and system based on knowledge transfer and attention mechanism. Background Art

[0002] Medical Image Segmentation (MIS) is a key technology in the fields of computer vision and medical image processing. Its goal is to separate different tissues and structures in medical images, so as to facilitate more in-depth analysis and research on specific regions in the images. This process involves identifying and labeling different parts in the image, such as organs, lesion areas, etc., so as to provide accurate image information for clinical diagnosis, treatment planning and medical research. Medical image segmentation plays a crucial role in the field of medical image processing, and it plays an indispensable role in computer-aided diagnosis, treatment planning and drug development. Therefore, the problem of medical image segmentation has received extensive attention.

[0003] However, medical image segmentation faces numerous challenges, which are mainly reflected in the following three aspects. First, the medical image dataset is often limited in size, especially the number of positive samples representing pathological or disease characteristics is small. Second, the annotation process of medical images is complex and requires the professional knowledge of medical professionals to accurately identify the affected areas, and the interpretation of these areas may vary due to differences in professional knowledge, experience or subjective judgment. Finally, medical image segmentation requires extremely high precision. To address these challenges, researchers have proposed medical image segmentation methods based on threshold-based methods, histogram-based segmentation, k-means clustering, watershed methods, active contours, conditional random fields, and ideas based on sparsity. However, these methods cannot accurately capture the essential characteristics of the region of interest, resulting in the need to improve the segmentation effect of the above methods. With the popularity of deep learning technology, more and more scholars have obtained more accurate segmentation results based on deep learning technology.

[0004] Existing medical image segmentation methods based on deep learning technology can be divided into two categories: CNN-based methods and Transformer-based methods. In CNN-based medical image segmentation methods, Yang Y and Feng C et al. proposed a medical image segmentation model that combines the deep learning model UNet with traditional image processing methods, and improves the accuracy of medical image segmentation by utilizing the complementary characteristics of the two methods in the model. Liu Y T and Zhu H J fused UNet and multi-layer perceptron structure in the article to construct Rolling-Unet and obtained relatively accurate medical image segmentation results, and verified the effectiveness of the proposed method on 4 public datasets. In Transformer-based medical image segmentation methods, Liu Z and Lin Y were among the first scholars to introduce Transformer technology into the field of computer vision. They introduced a windowed attention mechanism based on sliding windows, which reduces the computational complexity of the original ViT by calculating attention within the window and achieved excellent performance in multiple vision tasks. In methods that fuse CNN and Transformer, Cao H and Wang Y reconstructed the modules of Swin-Transformer according to the U-shaped structure and designed an encoder-decoder architecture with skip connections for medical image segmentation, and achieved excellent segmentation performance in multi-organ segmentation tasks. Zhang Y D and Liu H Y combined Transformer and CNN to propose a novel parallel branch architecture TransFuse. TransFuse can efficiently capture global dependencies and low-level spatial details and is implemented in a shallower manner.

[0005] Although existing medical image segmentation methods based on deep learning technology have achieved good results, CNN-based medical image segmentation methods have difficulties in learning clear long-range dependencies. Transformer-based medical image segmentation methods face problems such as high computational complexity and insufficient ability to learn local features. More importantly, there are problems such as blurred or discontinuous boundaries when marking segmentation boundaries in such methods.

[0006] Therefore, it is necessary to provide an improved technical solution to address the deficiencies of the above-mentioned existing technologies. Summary of the Invention

[0007] The purpose of this application is to provide a medical image segmentation method and system based on knowledge transfer and attention mechanism. Starting from the perspective of improving the applicability and robustness of existing medical image segmentation methods, based on the CNN structure, a module that can obtain long-range dependencies is constructed to obtain the mapping relationship between medical images and segmentation results, thereby effectively solving the problem of blurred boundaries in segmentation results.

[0008] To achieve the above object, the present application provides the following technical solutions:

[0009] The present application provides a medical image segmentation method based on knowledge transfer and attention mechanism, including:

[0010] Segmenting the medical image to be segmented based on the trained medical image segmentation model to obtain a segmentation result;

[0011] Among them, the medical image segmentation model includes a first sub-network and a second sub-network;

[0012] The first sub-network and the second sub-network adopt the network structure of an encoder and a decoder with the same structure; the encoder includes a plurality of first attention modules, the decoder includes a plurality of second attention modules with the same number as the encoder, and each first attention module of the encoder is jump-connected to the second attention module at the corresponding position of the decoder;

[0013] A knowledge transfer layer is provided between the first attention modules of the first sub-network and the first attention modules at the same position of the second sub-network, for transferring the first sub-features output by each first attention module of the first sub-network as knowledge to the second sub-network and performing feature fusion with the second sub-features, where the second sub-features are the sub-features output by each first attention module in the second sub-network, and the first sub-features correspond to the second sub-features.

[0014] In a possible implementation manner, in the first sub-network and the second sub-network, a plurality of consecutive Lo2 modules are provided between the encoder and the decoder.

[0015] In a possible implementation manner, before segmenting the medical image to be segmented based on the trained medical image segmentation model to obtain a segmentation result, it further includes:

[0016] Constructing a medical image segmentation model;

[0017] Training the medical image segmentation model using the training data set of medical images to obtain a trained medical image segmentation model.

[0018] In a possible implementation manner, training the medical image segmentation model using the training data set of medical images to obtain a trained medical image segmentation model includes:

[0019] Obtaining the training data set of medical images from multiple public data sets;

[0020] Inputting the label data in the training data set into the first sub-network to train the first sub-network;

[0021] Input the data to be segmented in the training dataset into the second sub-network and train the second sub-network.

[0022] In a possible implementation, during the training process of the medical image segmentation model, it further includes: calculating the total model loss;

[0023] The total model loss is the weighted sum of the first loss and the second loss;

[0024] The first loss is the sum of the cross-entropy losses between the first sub-features output by each feature extraction module of the first sub-network and the second sub-features output by the corresponding feature extraction modules at the same positions in the second sub-network;

[0025] The second loss is the DICE loss between the input and output of the second sub-network.

[0026] In a possible implementation, the first attention module and the second attention module have the same structure, both being multi-scale spatial pixel attention (MSPA) modules;

[0027] One average pooling layer follows each MSPA module in the first sub-network to form multiple feature extraction layers, and one bilinear interpolation layer follows each MSPA module in the second sub-network to form multiple feature restoration layers.

[0028] In a possible implementation, each MSPA module includes: a feature grouping unit, a multi-scale feature acquisition unit, a feature weight calculation unit, a feature normalization unit, a feature enhancement unit, and a feature activation unit;

[0029] The feature grouping unit is used to input the input feature into grouped convolution to obtain feature map a, feature map b, and feature map c;

[0030] The multi-scale feature acquisition unit is used to input feature map a and feature map b into two average pooling layers respectively, then perform 1×1 convolution and feature fusion on the output features of the two average pooling layers to obtain feature map d;

[0031] The feature weight calculation unit is used to input feature map d into the Sigmoid and Tanh operation functions respectively, and perform feature weight calculation on the output results of the Sigmoid and Tanh operations and feature map c to obtain feature map e;

[0032] The feature normalization unit is used to perform 2 times of 3×3 convolution and average pooling operations on feature map e, and then perform 1×1 convolution and batch normalization operations on the result of the average pooling operation to obtain feature map f;

[0033] The feature enhancement unit is used to perform average pooling and Softmax processing on feature map c to obtain feature map g;

[0034] A feature activation unit is configured to sequentially perform feature weight calculation and Softmax processing on the feature map g and the feature map f to obtain an output feature map.

[0035] Another embodiment of the present application provides a medical image segmentation system based on knowledge transfer and attention mechanism, including:

[0036] A segmentation unit is configured to segment the medical image to be segmented based on the trained medical image segmentation model to obtain a segmentation result;

[0037] Wherein, the medical image segmentation model includes a first sub-network and a second sub-network;

[0038] The first sub-network and the second sub-network adopt the network structures of an encoder and a decoder with the same structure; the encoder includes a plurality of first attention modules, the decoder includes a plurality of second attention modules with the same number as the encoder, and each first attention module of the encoder is jump-connected to the second attention module at the corresponding position of the decoder;

[0039] A knowledge transfer layer is provided between the first attention modules of the first sub-network and the first attention modules at the same position of the second sub-network, configured to transfer the first sub-features output by each first attention module of the first sub-network as knowledge to the second sub-network and perform feature fusion with the second sub-features, where the second sub-features are the sub-features output by each first attention module in the second sub-network, and the first sub-features correspond to the second sub-features.

[0040] The technical solution of the embodiment of the present application has the following beneficial effects:

[0041] In the technical solution of this application, the medical image segmentation model includes a first sub-network and a second sub-network; the first sub-network and the second sub-network adopt the network structures of an encoder and a decoder with the same structure; the encoder includes a plurality of first attention modules, and the decoder includes a plurality of second attention modules with the same number as the encoder. Through the dual sub-network structure and the stacked attention mechanism, the attention mechanism is integrated into the medical image segmentation task, so as to handle image information of multiple scales, effectively extract the non-linear expression ability of sub-features, and obtain fine-grained local features. Each first attention module of the encoder is skip-connected to the second attention module at the corresponding position of the decoder, which helps to restore detailed information and improve the accuracy of segmentation. A knowledge transfer layer is arranged between the first attention modules of the first sub-network and the first attention modules at the same position of the second sub-network, which is used to transfer the first sub-features output by each first attention module of the first sub-network as knowledge to the second sub-network and perform feature fusion with the second sub-features. Among them, the second sub-features are the sub-features output by each first attention module in the second sub-network, and the first sub-features correspond to the second sub-features. By integrating the knowledge transfer idea into the medical image segmentation task and mapping the sub-features as knowledge, the network structure based on CNN can also obtain long-range dependencies and effectively alleviate the problem of discontinuous boundaries in existing methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 FIG. is a schematic structural diagram of a medical image segmentation model (KT-UNet) provided according to some embodiments of the present application.

[0043] Figure 2 FIG. is a schematic diagram of the construction process of KT-UNet provided according to some embodiments of the present application.

[0044] Figure 3 FIG. is a schematic structural diagram of the MSPA module provided according to some embodiments of the present application.

[0045] Figure 4 FIG. is a schematic diagram of the medical image segmentation result obtained by using the method provided in this embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The terms "first", "second", "third", "fourth", etc. in the specification, claims and drawings of this application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.

[0047] As used herein, the mention of "embodiment" means that the specific features, structures, or characteristics described in connection with the embodiment may be included in at least one embodiment of the present application. The phrase appearing at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.

[0048] "Plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.

[0049] The embodiments of the present application will be described below with reference to the accompanying drawings.

[0050] Embodiment 1

[0051] This embodiment provides a medical image segmentation method based on knowledge transfer and attention mechanism, including:

[0052] Segmenting the medical image to be segmented based on the trained medical image segmentation model to obtain a segmentation result;

[0053] Among them, the medical image segmentation model includes a first sub-network and a second sub-network;

[0054] The first sub-network and the second sub-network adopt the network structures of the encoder and decoder with the same structure; the encoder includes a plurality of first attention modules, the decoder includes a plurality of second attention modules with the same number as the encoder, and each first attention module of the encoder is jump-connected to the second attention module at the corresponding position of the decoder;

[0055] A knowledge transfer layer is provided between the first attention modules of the first sub-network and the first attention modules at the same position of the second sub-network, for transferring the first sub-features output by each first attention module of the first sub-network as knowledge to the second sub-network and performing feature fusion with the second sub-features, where the second sub-features are the sub-features output by each first attention module in the second sub-network, and the first sub-features correspond to the second sub-features.

[0056] In this embodiment, the medical image segmentation model, also known as KT-UNet, refers to a machine learning or deep learning model used to identify and segment specific structures (such as organs, lesion regions) from medical images. Such as Figure 1As shown in the figure, KT-UNet consists of two sub-networks with the same structure, namely the first sub-network and the second sub-network. Among them, the first sub-network is also called the teacher network, and the second sub-network is also called the student network. Each sub-network has an encoder and a decoder, and both the encoder and the decoder have multiple attention modules. The attention module in the encoder is called the first attention module, which is used to perform successive downsampling on features. The attention module in the decoder is called the second attention module, which is used to perform successive upsampling operations on features. And each first attention module in the encoder has a skip connection with the second attention module at the corresponding position in the decoder. The setting of the skip connection helps to restore detailed information and improve the accuracy of segmentation. The use of multiple stacked attention modules is beneficial for the model to focus on important regions and improve the ability to capture local dependence relationships.

[0057] A knowledge transfer layer is set between the first attention modules at the same position between the two sub-networks, so that the first sub-features output by the first sub-network can be transferred to the second sub-network and fused with the corresponding sub-features of the second sub-network.

[0058] In this embodiment, knowledge is conceptualized as the mapping relationship between network inputs and their corresponding intermediate or final outputs. By using the intermediate or output data of the pre-trained network to guide the training of the new network, the transfer of the learned knowledge is realized. This knowledge transfer can effectively improve the generalization ability and robustness of the network.

[0059] In this embodiment, the trained medical image segmentation model refers to a deep learning (CNN) model trained by a large number of labeled medical image data (such as CT, MRI, etc.), which can automatically separate the target regions (such as organs, tumors) in the input medical image to be segmented and output the segmentation result.

[0060] In this embodiment, the medical image to be segmented is the medical image data that needs to extract the target region, usually an unlabeled or only partially labeled image.

[0061] In summary, for the method provided in this embodiment, the medical image segmentation model includes two sub-networks. Based on the CNN structure, from the perspective of improving the applicability and robustness of the existing medical image segmentation methods, a module that can obtain long-range dependence relationships is constructed by setting a knowledge transfer layer to obtain the mapping relationship between the medical image and the segmentation result, thereby effectively solving the problem of blurred boundaries of the segmentation result and improving the segmentation effect of the medical image.

[0062] In some embodiments, in the first sub-network and the second sub-network, multiple consecutive Lo2 modules are set between the encoder and the decoder. By using multiple Lo2 modules, the long-range dependence relationship of the medical image is obtained again, further improving the applicability and robustness of the model.

[0063] In some embodiments, before segmenting the medical image to be segmented based on the trained medical image segmentation model to obtain a segmentation result, the following steps are further included:

[0064] Construct a medical image segmentation model;

[0065] Use the training dataset of medical images to train the medical image segmentation model to obtain a trained medical image segmentation model.

[0066] In some embodiments, using the training dataset of medical images to train the medical image segmentation model to obtain a trained medical image segmentation model includes:

[0067] Obtain the training dataset of medical images from multiple public datasets;

[0068] Input the label data in the training dataset into the first sub-network to train the first sub-network;

[0069] Input the data to be segmented in the training dataset into the second sub-network to train the second sub-network.

[0070] In some embodiments, during the training process of the medical image segmentation model, the following steps are further included: calculating the total model loss; the total model loss is the weighted sum of the first loss and the second loss;

[0071] The first loss is the sum of the cross-entropy losses between the first sub-features output by each feature extraction module of the first sub-network and the second sub-features output by the corresponding feature extraction module of the second sub-network at the corresponding positions;

[0072] The second loss is the DICE loss between the input and output of the second sub-network.

[0073] By constructing a new loss function, the accuracy, robustness, and generalization ability of the proposed medical image segmentation method and system are improved.

[0074] In some embodiments, the first attention module and the second attention module have the same structure, both of which are multi-scale spatial pixel attention (MSPA) modules; after each MSPA module in the first sub-network, there is a following average pooling layer, forming multiple feature extraction layers, and after each MSPA module in the second sub-network, there is a following bilinear interpolation layer, forming multiple feature reduction layers.

[0075] Furthermore, each MSPA module includes: a feature grouping unit, a multi-scale feature acquisition unit, a feature weight calculation unit, a feature normalization unit, a feature enhancement unit, and a feature activation unit;

[0076] A feature grouping unit for inputting input features into grouped convolution to obtain feature map a, feature map b, and feature map c;

[0077] A multi-scale feature acquisition unit for inputting feature map a and feature map b into two average pooling layers respectively, then performing 1×1 convolution and feature fusion on the output features of the two average pooling layers to obtain feature map d;

[0078] A feature weight calculation unit for inputting feature map d into Sigmoid and Tanh operation functions respectively, and performing feature weight calculation on the output results of Sigmoid and Tanh operations and feature map c to obtain feature map e;

[0079] A feature normalization unit for performing 2 times of 3×3 convolution and average pooling operations on feature map e, and performing 1×1 convolution and batch normalization operations on the result of the average pooling operation to obtain feature map f;

[0080] A feature enhancement unit for performing average pooling and Softmax processing on feature map c to obtain feature map g;

[0081] A feature activation unit for sequentially performing feature weight calculation and Softmax processing on feature map g and feature map f to obtain an output feature map.

[0082] The technical solution of this embodiment will be illustrated by examples below.

[0083] As an example, there are 4 MSPA modules in both the encoder and the decoder, and the number of Lo2 modules is configured to be 3.

[0084] As Figure 2 shown, this embodiment constructs KT-UNet according to the following steps:

[0085] (1) Design a multi-scale pixel spatial attention (MSPA) module.

[0086] Spatial relationships ensure the relative positions and neighborhood relationships of pixels in medical images, providing important clues and fine-grained local features for image segmentation. In addition, pixel-level relationships are crucial for obtaining edge and boundary definitions as well as texture and pattern recognition. In addition, to address the segmentation effects of medical images at different scales, multi-scale relationships are introduced to capture multi-scale significant features. Based on this, this embodiment designs an MSPA module, the structure of which is as Figure 3 shown.

[0087] The construction of the MSPA module includes the following steps: First, analyze the essential reasons for the large computational complexity of existing methods and construct a feature grouping unit; Second, to obtain multi-scale features, construct a multi-scale feature acquisition unit; Third, to enhance the non-linear expression ability of sub-features and obtain fine-grained local features, a feature weight operation unit; To enhance the ability of sub-features to learn fine-grained information, construct a feature normalization unit; Third, to establish the association between the extracted features and the image segmentation problem, construct a feature enhancement unit; Finally, to obtain the operation relationship between fine-grained information and structured information, construct a feature activation unit.

[0088] Among them, to improve the computational efficiency, the feature grouping unit in the MSPA module is used to perform the following steps: The initial feature (i.e., the input feature) is defined as Input Into the grouped convolution, and the number of grouped convolutions is defined as G. The three types of features output by the grouped convolution are respectively defined as (i.e., feature map a), (i.e., feature map b) and (i.e., feature map c).

[0089] To obtain multi-scale features, the multi-scale feature acquisition unit in the MSPA module is used to perform the following steps: Input And Into two average pooling layers respectively. Then, the features output by the two average pooling layers are input into a 1×1 convolution and a feature fusion layer, and (i.e., feature map d) is obtained. The above process can be defined in the following form:

[0090]

[0091] Among them, f 1×1 () is a 1×1 convolution, Avg.Pool() is an average pooling operation, Is a feature fusion operation.

[0092] To enhance the non-linear expression ability of sub-features and obtain fine-grained local features, the feature weight operation unit in the MSPA module is used to perform the following steps: Are respectively input into the Sigmoid and Tanh operation functions, and then the outputs of the Sigmoid and Tanh operation units are used to perform a feature weight operation with To obtain (i.e., feature map e). The above process can be defined in the following form.

[0093]

[0094] Among them, Rewight() is the feature weight operation function, Sigmoid() is the Sigmoid operation function, and Tanh() is the Tanh operation function.

[0095] To enhance the ability of sub-features to learn fine-grained information, the feature specification unit in the MSPA module is used to perform the following steps: They are respectively input into two 3×3 convolutional layers and an average pooling layer, and then the average pooling layer is subjected to 1×1 convolution and batch normalization operations to obtain (i.e., the feature map f). The above process can be defined in the following form:

[0096]

[0097] Among them, BN() is the batch normalization operation, and f 3×3 () is the 3×3 convolution operation,

[0098] To establish the association between the extracted features and the image segmentation problem, the feature enhancement unit in the MSPA module is used to perform the following steps: It is input into the average pooling layer and the Softmax layer to obtain (i.e., the feature map g). The above process can be described in the following form.

[0099]

[0100] Among them, Softmax() is the Softmax operation.

[0101] To obtain the operation relationship between fine-grained information and structured information, the feature activation unit in the MSPA module is used to perform the following steps: and perform feature weight calculation and Softmax processing successively, and the output feature is the feature obtained by MSPA (i.e., the output feature map).

[0102] The above is an exemplary description of the various components of the MSPA module.

[0103] (2) On the basis of constructing the MSPA module, integrate the MSPA to construct the Knowledge Transfer UNet (KT_UNet) that combines the MSPA module and knowledge transfer, including:

[0104] First, to obtain local dependencies, the constructed MSPA is integrated; second, to improve the acquisition effect of long-range dependencies, the Lo2 module is introduced to construct the first sub-network; finally, to solve the problem of blurred boundaries in the segmentation results, based on knowledge transfer, the sub-features generated by MSPA are used as knowledge to complete the construction of KT-UNet.

[0105] The detailed description of each step is as follows:

[0106] To obtain local dependencies, first, the medical image is input into a 3×3 convolution, and then the output of the 3×3 convolution is continuously input into 4 MSPA modules, and each MSPA module is followed by an average pooling layer to achieve downsampling of the features.

[0107] Then, the following steps are performed to construct the first sub-network:

[0108] Among them, to further improve the extraction effect of long-range dependencies, the output of the average pooling layer after the fourth MSPA module is continuously input into 3 Lo2 modules; the output of the third Lo2 module is continuously input into 4 MSPA modules for decoding, and each MSPA module is followed by a bilinear interpolation to achieve upsampling of the features; again, the output of the bilinear interpolation after the eighth MSPA is input into a 3×3 convolution; finally, skip connections are made between the first MSPA and the eighth MSPA, the second MSPA and the seventh MSPA, the third MSPA and the sixth MSPA, and the fourth MSPA and the fifth MSPA to complete the construction of the first sub-network.

[0109] Finally, to solve the problem of blurred boundaries in the segmentation results, another sub-network with the same structure, i.e., the second sub-network, is constructed according to the construction process of the first sub-network; then, the first sub-network is used as the teacher network and the second sub-network is used as the student network, and the outputs of the MSPAs at the same positions in the sub-networks with the same structure are fused to complete the construction of KT-UNet.

[0110] (3) Construct a loss function based on the idea of knowledge transfer. Combining the public dataset, the obtained knowledge is converted into a loss to construct a loss function suitable for KT-UNet.

[0111] First, select public datasets such as BKAI, ClinicDB, ColonDB, Kvasir, ISIC2017, ISIC2018, BMoNuSeg, and CHASEDB1 as the training datasets for the embodiments of this application; then, input the GroundTruth (i.e., label data) in the above datasets into the teacher network to implement the training of the teacher network; subsequently, input the images to be segmented in the above datasets into the student network, calculate the cross-entropy loss corresponding to the 8 MSPA output features corresponding in the teacher network and the student network, and denote it as l KT ; finally, calculate the DICE loss between the input and output of the student network and denote it as l cd , and the total loss of the model is as follows;

[0112]

[0113] where l is the overall loss of KT-UNet (i.e., the total loss of the model), α and β are balance coefficients, M is the number of MSPA modules in the teacher network or the student network, i is the MSPA index, is the output of the i-th MSPA in the teacher network, is the output of the i-th MSPA in the student network. The training of KT-UNet is implemented through the selected datasets and loss functions, and the segmentation of images such as polyps, cells, and blood vessels is achieved.

[0114] In this embodiment, KT-UNet includes a first sub-network (teacher network) and a second sub-network (student network), and the teacher network can be used to guide the student network, thereby improving the generalization ability and robustness of the medical image segmentation model. By constructing the loss function with the obtained knowledge, l KT represents the knowledge transfer loss, l cd represents the DICE loss between the input and output of the student network. Through weighted addition by the balance coefficients α and β, the two losses are combined to achieve collaborative optimization, enabling the model to generate more accurate segmentation results while maintaining efficient feature extraction.

[0115] It should be noted that for the medical image segmentation model provided in this embodiment, although its structure involves a teacher network and a student network, the basic idea of this model is knowledge transfer rather than knowledge distillation. The set knowledge transfer layer can map each sub-feature of the first sub-network as knowledge into the second sub-network, so that the second sub-network can obtain the long-distance dependence relationship between medical images. During the training process, l KT is added to the loss function to convert the transferred knowledge into loss, thereby improving the accuracy, robustness, and generalization ability of the proposed medical image segmentation method and system.

[0116] Figure 4shows the medical image segmentation results obtained by using the method provided in this embodiment. From Figure 4 it can be seen that the method provided in this embodiment can accurately segment the target region in the medical image, and the boundary is clear and complete.

[0117] Thus, it can be seen that the medical image segmentation method based on knowledge transfer and attention mechanism provided in this embodiment can obtain multi-scale features of the image to be segmented through the multi-scale spatial pixel attention mechanism, enhance the non-linear expression ability of sub-features, and obtain fine-grained local features. By constructing two sub-networks with the same structure and performing knowledge transfer between the two sub-networks, it can effectively obtain local dependence relationships, enhance the boundary information of the image to be segmented, and alleviate the boundary blurring problem of medical image segmentation in the prior art; construct a loss function based on the idea of knowledge transfer to obtain the mapping relationship between the image to be segmented and the segmentation result, thereby improving the segmentation effect of medical images.

[0118] Embodiment 2:

[0119] Based on the same inventive concept, an embodiment of the present application provides a medical image segmentation system based on knowledge transfer and attention mechanism, including:

[0120] A segmentation unit for segmenting the medical image to be segmented based on the trained medical image segmentation model to obtain a segmentation result;

[0121] Wherein, the medical image segmentation model includes a first sub-network and a second sub-network;

[0122] The first sub-network and the second sub-network adopt a network structure of an encoder and a decoder with the same structure; the encoder includes a plurality of first attention modules, the decoder includes a plurality of second attention modules with the same number as the encoder, and each first attention module of the encoder is jump-connected to the second attention module at the corresponding position of the decoder;

[0123] A knowledge transfer layer is provided between the first attention modules of the first sub-network and the first attention modules at the same position of the second sub-network, for transferring the first sub-features output by each first attention module of the first sub-network as knowledge to the second sub-network and performing feature fusion with the second sub-features, where the second sub-features are the sub-features output by each first attention module in the second sub-network, and the first sub-features correspond to the second sub-features.

[0124] The medical image segmentation system based on knowledge transfer and attention mechanism provided in this embodiment can implement the processes and steps of the medical image segmentation method based on knowledge transfer and attention mechanism provided in any of the above embodiments, and achieve the same technical effects, which will not be elaborated here one by one.

[0125] The foregoing are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and variations can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A medical image segmentation method based on knowledge transfer and attention mechanism, characterized in that Including: Segmenting the medical image to be segmented based on the trained medical image segmentation model to obtain a segmentation result; wherein, the medical image segmentation model includes a first sub-network and a second sub-network; The first sub-network and the second sub-network adopt the network structure of an encoder and a decoder with the same structure; the encoder includes a plurality of first attention modules, the decoder includes a plurality of second attention modules with the same number as the encoder, and each first attention module of the encoder is skip-connected to the second attention module at the corresponding position of the decoder; A knowledge transfer layer is arranged between the first attention modules of the first sub-network and the first attention modules at the same position of the second sub-network, which is used to transfer the first sub-features output by each first attention module of the first sub-network as knowledge to the second sub-network and perform feature fusion with the second sub-features, wherein the second sub-features are the sub-features output by each first attention module in the second sub-network, and the first sub-features correspond to the second sub-features.

2. The method according to claim 1, characterized in that, In both the first sub-network and the second sub-network, a plurality of consecutive Lo2 modules are arranged between the encoder and the decoder.

3. The method according to claim 1, characterized in that Before segmenting the medical image to be segmented based on the trained medical image segmentation model to obtain a segmentation result, it further includes: Constructing a medical image segmentation model; Using the training data set of medical images to train the medical image segmentation model to obtain a trained medical image segmentation model.

4. The method according to claim 3, characterized in that, Using the training data set of medical images to train the medical image segmentation model to obtain a trained medical image segmentation model, including: Obtaining the training data set of medical images from multiple public data sets; Inputting the label data in the training data set into the first sub-network to train the first sub-network; Inputting the data to be segmented in the training data set into the second sub-network to train the second sub-network.

5. The method according to claim 4, wherein During the training process of the medical image segmentation model, it further includes: calculating the total model loss; The total model loss is the weighted sum of the first loss and the second loss; The first loss is the sum of the cross-entropy losses between the first sub-features output by each feature extraction module of the first sub-network and the second sub-features output by the feature extraction modules at the corresponding positions of the second sub-network; The second loss is the DICE loss between the input and output of the second sub-network.

6. The method according to claim 1, wherein The first attention module and the second attention module have the same structure, both of which are multi-scale spatial pixel attention (MSPA) modules; One average pooling layer follows each MSPA module in the first sub-network to form a plurality of feature extraction layers, and one bilinear interpolation layer follows each MSPA module in the second sub-network to form a plurality of feature restoration layers.

7. The method according to claim 6, wherein Each MSPA module includes: a feature grouping unit, a multi-scale feature acquisition unit, a feature weight operation unit, a feature normalization unit, a feature enhancement unit, and a feature activation unit; The feature grouping unit is configured to input the input feature into grouped convolution to obtain feature map a, feature map b, and feature map c; A multi-scale feature acquisition unit for inputting feature map a and feature map b into two average pooling layers respectively, then performing 1×1 convolution and feature fusion on the output features of the two average pooling layers to obtain feature map d; A feature weight calculation unit for inputting feature map d into Sigmoid and Tanh operation functions respectively, and performing feature weight calculation on the output results of the Sigmoid and Tanh operations and feature map c to obtain feature map e; A feature normalization unit for performing 2 times of 3×3 convolution and average pooling operations on feature map e, and performing 1×1 convolution and batch normalization operations on the results of the average pooling operation to obtain feature map f; A feature enhancement unit for performing average pooling and Softmax processing on feature map c to obtain feature map g; A feature activation unit for sequentially performing feature weight calculation and Softmax processing on feature map g and feature map f to obtain an output feature map.

8. A medical image segmentation system based on knowledge transfer and attention mechanism, comprising: A segmentation unit for segmenting a medical image to be segmented based on a trained medical image segmentation model to obtain a segmentation result; Wherein, the medical image segmentation model includes a first sub-network and a second sub-network; The first sub-network and the second sub-network adopt the network structures of an encoder and a decoder with the same structure; the encoder includes a plurality of first attention modules, the decoder includes a plurality of second attention modules with the same number as the encoder, and each first attention module of the encoder is skip-connected to the second attention module at the corresponding position of the decoder; A knowledge transfer layer is arranged between the first attention modules of the first sub-network and the first attention modules at the same positions of the second sub-network for transferring the first sub-features output by each first attention module of the first sub-network as knowledge to the second sub-network and performing feature fusion with the second sub-features, wherein the second sub-features are the sub-features output by each first attention module in the second sub-network, and the first sub-features correspond to the second sub-features.

Citation Information

Cited By

  • Medical image segmentation method, system and device

    CN121259323A

  • Medical record AI intelligent integration and analysis system based on big data

    CN121545656A