Knee joint MRI tibiofemoral joint tissue segmentation method based on multi-scale feature fusion
Through the improved VM-Unet architecture and multi-scale fusion network, the problems of insufficient fusion of multi-scale features and blurred boundaries in knee MRI image segmentation are solved, and high-precision knee tissue segmentation is achieved, supporting early diagnosis and personalized treatment.
Patent Information
- Application Number
- CN202510443155.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-11
AI Technical Summary
In the MRI image segmentation of knee joint, it is difficult to take into account the detailed characteristics of the microcartilage structure and the global anatomical structure characteristics of the entire knee joint in the prior art. The global context modeling efficiency is low and the boundary segmentation accuracy is insufficient, which affects the accurate diagnosis and treatment of knee joint diseases.
Using the improved VM-Unet architecture, a multi-scale fusion network (MSPF-VM-Unet) is used to achieve high-precision and efficient knee tissue segmentation through the multi-scale pyramid feature extraction network (MPSK-Net) and Selective Kernel (SK) attention mechanism, combined with the Efficient Channel Attention (ECA) mechanism.
The segmentation accuracy of joint tissue in knee MRI images is improved, technical support for early diagnosis and personalized treatment is provided, and the problems of insufficient fusion of multi-scale features and blurred boundaries in traditional models in knee segmentation are solved.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and particularly to a method for segmenting tibiofemoral joint tissues in knee MRI based on multi-scale feature fusion, which is used to accurately segment joint tissues from knee MRI images to assist in the diagnosis and treatment of knee diseases. Background Art
[0002] Knee arthritis, as a global chronic disease, has an incidence rate as high as 22.9% in people over 40 years old, and the incidence rate shows an upward trend with the increase of obesity rate and age. The accurate diagnosis and treatment planning of knee joint diseases such as osteoarthritis highly depend on the fine segmentation of knee joint tissues such as cartilage and meniscus.
[0003] With the rapid development of medical imaging technologies such as computed tomography (CT) and magnetic resonance imaging (MRI), medical image segmentation methods based on deep learning have become the core tools for clinical auxiliary diagnosis. However, due to the complex anatomical structure, blurred tissue boundaries, and susceptibility to imaging noise interference of knee joint tissues, many challenges are brought to traditional segmentation models. For example, traditional models often have problems of insufficient multi-scale feature fusion, making it difficult to simultaneously consider the detailed features of tiny cartilage structures and the global anatomical structure features of the entire knee joint; the global context modeling efficiency is low, and the ability to handle long-range dependence relationships is limited, such as performing poorly in associating the feature of cartilage tissues at different levels; the boundary segmentation accuracy is limited, resulting in blurred boundaries of the segmentation results and affecting the accurate diagnosis and treatment of knee joint diseases.
[0004] Although deep learning has gradually taken the leading position in the field of image segmentation, and some networks and their variants have made remarkable progress in medical image segmentation, there are still problems such as local detail loss and blurred segmentation effect boundaries when dealing with the fine-grained segmentation tasks in the multi-tissue intersection areas of the knee joint. Therefore, it is of great clinical significance to develop a knee MRI image segmentation method that can effectively solve the above problems. Summary of the Invention
[0005] The present invention aims to propose a multi-scale fusion network (MSPF-VM-Unet) based on an improved VM-Unet architecture, achieve a balance between high precision and high efficiency through innovative design, improve the segmentation accuracy of joint tissues in knee MRI images, and provide reliable technical support for the early diagnosis and personalized treatment planning of knee diseases.
[0006] To achieve the above object, the technical solution of the present invention is as follows: A method for segmenting tibiofemoral joint tissues in knee MRI based on multi-scale feature fusion, comprising the following steps: Step 1: Preprocess the data used, specifically including the following sub-steps: (a) Split the MRI sequence into individual slices and convert the slice format; (b) Process the MRI images to enhance the image contrast; (c) Perform min-max normalization on the image data of the MRI sequence and augment the training set samples; Step 2: Construction of the Multi-Scale Pyramid Feature Extraction Network (MPSK-Net), which specifically includes the following sub-steps: (a) Improve the traditional convolutional neural network and propose the Multi-Scale Pyramid Feature Extraction Network; (b) Replace the standard convolutional kernels in the traditional residual network with the Multi-Scale Pyramid Attention (MSPA) module combined with the Selective Kernel (SK) attention mechanism, and remove the global average pooling layer and the fully connected layer to meet the requirements of the segmentation task for high-resolution feature maps; (c) By parallelly using convolutional kernels of different sizes (such as 1×1, 3×3, 5×5) and the pyramid pooling module, MPSK-Net can accurately capture features at different scales and achieve efficient local and global information fusion; Step 3: Extract features using the parallel feature extraction network, which specifically includes the following sub-steps: (a) Run the Multi-Scale Pyramid Feature Extraction Network (MPSK-Net) in parallel with the VM-Unet encoder to extract multi-scale local features and global context features respectively; perform feature fusion and send it to the encoder; (b) Fuse the concatenated features and send them to the encoder; Step 4: Connect the output features to the decoder through the skip connection layer, which specifically includes the following sub-steps: (a) Introduce the Efficient Channel Attention (ECA) mechanism at the skip link to optimize the feature fusion mechanism of the skip connection; (b) Calculate the adaptive weights between channels using 1D convolution, and finally apply the attention weights to the original feature map for skip connection fusion with the decoder features; Step 5: Restore the output feature map to the corresponding label map through the decoder. Description of the Drawings
[0007] Figure 1 This is the main framework diagram of the present invention; Figure 2 This is the structural diagram of the Multi-Scale Pyramid Feature Extraction Network (MPSK-Net) of the present invention; Figure 3This is the structural diagram of the multi-scale pyramid feature extraction module (MPSK Block) of the present invention; Detailed implementation manners
[0008] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted. The technical solutions of the present invention will be further described below with reference to the drawings and embodiments.
[0009] As Figure 1 shown, a segmentation method for knee MRI images based on multi-scale feature fusion includes the following steps: Step 1: Preprocess the data used; Step 1.1: Convert the DICOM format MRI sequence to the NIfTI standard format, and crop the sequence to 160×160 pixels; Step 1.2: Remove the starting and ending slices in each MRI sequence. The removed slices do not contain bone or cartilage information and also contain excessive noise; Step 1.3: Use the CLAHE technique to enhance the contrast of the MRI image, with the parameter settings of clip limit = 40 and tile grid = 8×8; perform data augmentation through random rotation (±15°), elastic deformation (α = 300, σ = 20), and gamma correction (γ = 0.8 - 1.2).
[0010] Step 2: Construct a multi-scale pyramid feature extraction network (MPSK-Net); Step 2.1: As Figure 2 shown, the extraction network for MRI local features is composed of a standard convolutional block and a multi-scale pyramid feature extraction module (MPSK Block), and feature extraction is performed through residual connection; Step 2.2: As Figure 3 shown, the MPSK Block is composed of a pyramid pooling and an SK attention mechanism, and the extraction of local shallow, middle, and high-level features is controlled by 1×1, 2×2, and 4×4 convolutional kernels. The extracted features will be weighted in channels through the SK attention mechanism; Step 2.3: Concatenate the multi-scale features extracted in the above steps; Step 3: As Figure 1 shown, fuse the multi-scale local features extracted by MPSK-Net with the global features extracted by VM-Unet, and send them to the next layer to continue extracting and fusing multi-scale features to better extract features and achieve the complementary advantages of local multi-scale information and global multi-scale features; Step 4: AsFigure 1 As shown, the features output by the multi-scale feature fusion network module are fed into the ECA module; Step 4.1: Calculate the channel weights through the ECA module for feature enhancement; Step 4.2: Feed the features enhanced by the ECA module into the decoder.
[0011] Step 5: Restore the output fused features to the corresponding segmentation result map through the decoder; Step 5.1: As Figure 1 shown, a five-layer decoder is built on the basis of the VM-Unet architecture, and the corresponding segmentation result map is restored without reducing the receptive field by stacking transposed convolutions with a stride of 2 and a padding of 1 for 4×4, and ordinary convolutional kernels with a size of 3×3, a stride of 2, and a padding of 1.
Claims
1. A method for segmenting the tibiofemoral joint tissues in knee MRI based on multi-scale feature fusion, characterized in that, This method integrates medical image information of multiple scales to extract the features of knee MRI images and segment tissues. First, a new multi-scale pyramid feature extraction network MPSK-Net is designed. MPSK-Net enhances multi-scale local feature extraction through selective convolution (SK) and pyramid pooling. Then, the features extracted by MPSK-Net are incorporated into the VM-Unet architecture in parallel, enabling the model to capture the global information of MRI knee tissues while maintaining the ability to extract multi-scale local features. The fused features are linked to each layer of the decoder through a channel attention (ECA) module to improve the segmentation ability of small target tissues. Finally, the decoder part restores the feature map and obtains the final segmentation result. This design can capture local and global information of multiple scales and effectively fuse the two, improving the segmentation accuracy of knee joint tissues in MRI, thus enhancing the accuracy of computer-aided diagnosis. The specific steps are as follows: Step 1: Preprocess the dataset; Step 1.1: Convert the DICOM-format MRI sequence to the NIfTI standard format and crop the sequence to 160×160 pixels; Step 1.2: Remove the starting and ending slices in each MRI sequence. The removed slices do not contain bone or cartilage information and also contain excessive noise; Step 1.3: Use the CLAHE technique to enhance the contrast of the MRI images, with the parameter settings of clip limit = 40 and tilegrid = 8×8; perform data augmentation through random rotation (±15°), elastic deformation (α = 300, σ = 20), and gamma correction (γ = 0.8 - 1.2); Step 2: Construct the multi-scale pyramid feature extraction network MPSK-Net; Step 2.1: As shown in Figure 2, the extraction network of MRI local features consists of a standard convolutional block and a multi-scale pyramid feature extraction module (MPSK Block), and feature extraction is carried out through residual connection; Step 2.2: As shown in Figure 3, the MPSK Block consists of pyramid pooling and the SK attention mechanism, and the extraction of local shallow, middle, and high-level features is controlled by convolutional kernels of 1×1, 2×2, and 4×4. The extracted features will be weighted in the channel through the SK attention mechanism; Step 2.3: Concatenate the multi-scale features extracted in the above steps; Step 3: As shown in Figure 1, fuse the multi-scale local features extracted by MPSK-Net and the global features extracted by VM-Unet, send them to the next layer to continue extracting and fusing multi-scale features to more comprehensively extract features, and achieve the complementary advantages of local multi-scale information and global multi-scale features; Step 4: As shown in Figure 1, send the features output by the multi-scale feature fusion network module to the ECA module; Step 4.1: Calculate the channel weights through the ECA module for feature enhancement; Step 4.2: Feed the features enhanced by the ECA module into the decoder; Step 5: Restore the output fused features to the corresponding segmentation result map through the decoder; Step 5.1: As shown in Figure 1, a five-layer decoder is built on the basis of the VM-Unet architecture. By stacking transposed convolutions with a stride of 2 and a padding of 1 for 4×4, and ordinary convolution kernels with a size of 3×3, a stride of 2, and a padding of 1, the corresponding segmentation result map is restored without reducing the receptive field.