Multi-direction dual-module medical image segmentation method and system based on MRI (Magnetic Resonance Imaging) image

By combining a dual-module structure and an innovative loss function, efficient and accurate segmentation of anal fistula MRI images is achieved, solving the problems of wasted computing resources and low segmentation accuracy in existing technologies, and making it suitable for clinical auxiliary diagnosis in primary hospitals.

CN121391902APending Publication Date: 2026-01-23ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511487696.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing MRI image segmentation techniques for anal fistula detection suffer from problems such as wasted computational resources, low segmentation accuracy, and poor interpretability. In particular, when experienced doctors are lacking in primary hospitals, the misdiagnosis rate and recurrence rate are high.

Method used

A dual-module medical image segmentation method based on multi-directional MRI images is adopted, including a preliminary segmentation and localization module and a multi-dimensional refinement module. It combines the Dice loss function with residual connectivity and outlier penalty to achieve efficient fusion and fine segmentation of multi-modal and multi-directional information.

Benefits of technology

It improves segmentation accuracy and model interpretability, reduces computational resource consumption, significantly reduces false positive rate, and is suitable for clinical application in primary hospitals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121391902A_ABST
    Figure CN121391902A_ABST
Patent Text Reader

Abstract

The invention provides a multi-direction dual-module medical image segmentation method and system based on an MRI (Magnetic Resonance Imaging) image. The method comprises the following steps: constructing an image segmentation model comprising a preliminary segmentation and positioning module and a multi-dimensional refinement module; the method comprises the following steps: preprocessing original multidirectional MRI image data, determining a main direction according to a task requirement, and obtaining a two-dimensional slice of the main direction; the two-dimensional slice in the main direction is input into a preliminary segmentation and positioning module to perform preliminary segmentation and positioning on the focus, and a preliminary segmentation probability graph is output; according to the preliminary segmentation probability graph, obtaining a slice highly related to the spatial position of the segmentation area in the main direction in each of other directions, and splicing the slices with the two-dimensional slices in the main direction to obtain a multi-channel composite feature; the multi-channel composite features are input into a multi-dimensional refining module for deep fusion and fine segmentation, and fine segmentation of the focus is completed. The invention provides a novel double-module structure, multi-mode and multi-direction information can be efficiently fused, and segmentation precision and model interpretability are both considered.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of medical image segmentation, and particularly relates to a multi-directional dual-module medical image segmentation method and system based on MRI images. BACKGROUND

[0002] In recent years, deep learning technology has made significant progress in the field of medical image segmentation, especially in the segmentation of human organs and tissues. The accuracy of medical image segmentation is of great significance for accurate diagnosis of diseases and subsequent treatment plan. Taking the magnetic resonance image of anal fistula as an example, due to the high similarity of lesions and normal tissues in the image, it is easy to have a high false positive rate in the segmentation process, which brings great challenges to clinical diagnosis.

[0003] Anal fistula is a common perianal disease, characterized by irregular abnormal channels between perianal skin and anal canal. Because of the high similarity between anal fistula and anal canal structure, and the possible intersection or connection between them under different conditions, doctors face great difficulty in interpreting MRI images. Especially in primary hospitals, although they have MRI imaging equipment, they lack experienced radiologists or surgeons, which can easily lead to misdiagnosis or missed diagnosis, and thus affect the treatment effect and recurrence rate.

[0004] MRI technology can provide multi-directional image information such as coronal, sagittal and transverse, which is an important tool for comprehensive assessment and surgical planning of anal fistula. However, doctors need to have high professional knowledge to fully utilize multi-modal MRI information. In recent years, the combination of deep learning and multi-modal MRI information has provided doctors with auxiliary diagnosis suggestions, effectively improving the accuracy and efficiency of MRI image segmentation.

[0005] Currently, some studies have tried to use multi-modal MRI features to improve the segmentation accuracy of different tissue types. For example, by fusing T1, T2 and other different sequence parameters, the visibility between tissues is enhanced, and the accuracy of lesion positioning and segmentation is improved. In addition, some studies have proposed multi-modal feature fusion models that can process multi-sequence MRI images and extract key information to improve segmentation results. Some scholars use multi-directional MRI images, model each sequence respectively, and integrate the segmentation results of each model through an integrated evaluation mechanism. With the improvement of computing power, three-dimensional convolutional neural networks have also been applied to the segmentation of three-dimensional MRI images. However, existing researches mostly focus on brain and chest image segmentation, and there are relatively few studies on pelvic and anal fistula lesions.

[0006] Despite advancements in MRI image segmentation, numerous challenges remain in the application of 3D fusion techniques for anal fistula detection. Current fusion strategies primarily fall into two categories: shallow fusion at the model's input and output, and deep fusion within the model itself. Shallow fusion requires the model to process the entire image sequence, leading to the handling of a large amount of redundant information. This wastes computational resources and results in insufficient expressive power during the encoding stage and low feature relevance during the decoding stage. While deep fusion improves feature expressive power, it suffers from a large number of model parameters, making it prone to overfitting and resulting in high training costs. These issues limit the full utilization of the multimodal information advantages of MRI and also affect the generalization ability and practical application value of segmentation models.

[0007] When processing multimodal and multi-directional MRI images, existing methods typically require processing large amounts of data, resulting in low efficiency in model training and inference. For lesions with complex structures and blurred boundaries, such as anal fistulas, existing segmentation methods suffer from low accuracy and poor interpretability in terms of localization and fine segmentation. Summary of the Invention

[0008] To address the aforementioned issues, this invention proposes a dual-module medical image segmentation method and system based on multi-directional MRI images, which can efficiently fuse multi-modal and multi-directional information while balancing segmentation accuracy and model interpretability.

[0009] To achieve the above objectives, the present invention provides a dual-module medical image segmentation method based on multi-directional MRI images, comprising:

[0010] An image segmentation model is constructed, which includes a preliminary segmentation and localization module and a multi-dimensional refinement module;

[0011] The raw multi-directional MRI image data is preprocessed, and the main direction is determined and two-dimensional slices of the main direction are obtained according to the task requirements.

[0012] The two-dimensional slices in the main direction are input into the preliminary segmentation and localization module to perform preliminary segmentation and localization of the lesions, and output a preliminary segmentation probability map;

[0013] Based on the preliminary segmentation probability map, obtain a slice in each of the other directions that is highly correlated with the spatial position of the segmented region in the main direction, and stitch it with the two-dimensional slice in the main direction to obtain a multi-channel composite feature.

[0014] The multi-channel composite features are input into the multi-dimensional refinement module for deep fusion and fine segmentation, thereby completing the fine segmentation of the lesions.

[0015] This invention proposes a novel dual-module structure, including a preliminary segmentation and localization module and a multidimensional refinement module. This structure achieves coarse localization of lesions through preliminary segmentation, and on this basis, utilizes the multidimensional refinement module to perform deep fusion and fine segmentation of multimodal and multi-directional MRI information, resulting in high segmentation accuracy and good model interpretability.

[0016] As a preferred embodiment, a dual-module medical image segmentation method based on multi-directional MRI images further includes:

[0017] After being activated by Sigmoid, the preliminary segmentation probability map is input as the residual term to the end of the multidimensional refinement module. It is then summed pixel by pixel with the output of the main branch of the multidimensional refinement module to obtain the final segmentation probability map.

[0018] By enabling direct information transfer between the preliminary segmentation and localization module and the multidimensional refined segmentation module through residual terms, the multidimensional refined segmentation module can further refine and correct the segmentation results based on the preliminary localization by the preliminary segmentation and localization module.

[0019] As a preferred embodiment, a dual-module medical image segmentation method based on multi-directional MRI images further includes:

[0020] The model is trained using the Dice loss function, which introduces outlier penalties, to penalize segmentation edges and outliers.

[0021] To reduce the false positive rate, the Dice loss function (OPDL) with outlier penalty is introduced to penalize segmentation edges and outliers, which effectively improves the accuracy of segmentation boundaries and overall performance.

[0022] Preferably, the preprocessing of the original multi-directional MRI image data, determining the principal direction according to task requirements, and obtaining two-dimensional slices of the principal direction includes the following steps:

[0023] The original multi-directional MRI image data was parsed in a unified format to extract the coronal, sagittal and transverse image data, and then converted into standard three-dimensional volume data.

[0024] Based on the spatial distribution characteristics of anal fistula lesions and the clinical diagnostic process, the coronal plane was selected as the main direction.

[0025] The process iterates through the 3D volume data, slices it along the main direction to generate a sequence of 2D slices, and records the spatial index of each 2D slice in the 3D volume data. This facilitates subsequent spatial alignment and multi-directional information fusion.

[0026] Preferably, the preprocessing of the original multi-directional MRI image data, determining the principal direction according to task requirements, and obtaining two-dimensional slices of the principal direction further includes the following steps:

[0027] To eliminate resolution and size differences caused by different patients and different scanning batches, all two-dimensional slices must undergo size normalization and resampling.

[0028] To further eliminate inconsistencies in grayscale distribution caused by differences in equipment and batches, the system performs grayscale normalization processing on the two-dimensional slices after size normalization and resampling.

[0029] The corresponding labels are obtained by synchronously loading the manually annotated lesion segmentation mask corresponding to each two-dimensional slice.

[0030] Preferably, the step of obtaining a slice in each of the other directions based on the preliminary segmentation probability map that is highly correlated with the spatial position of the segmented region in the main direction, and stitching it with the two-dimensional slice in the main direction to obtain a multi-channel composite feature includes the following steps:

[0031] The preliminary segmentation probability map is analyzed at the pixel level, and a set of all pixels with a foreground probability greater than a set threshold is selected. To avoid low-confidence pixels interfering with the spatial center point, the system adjusts the probability value w before calculation. i Threshold truncation is performed, retaining only pixels above the set threshold for calculation; preferably, if the current slice is a lesion-free or healthy sample, the overall probability output by the preliminary segmentation and localization module is low, the system automatically sets the center point to the image center (256, 256) and skips the iterative update.

[0032] For each selected pixel, record its spatial coordinates (x, y) in the corresponding 2D slice. i ,y i ) and its corresponding probability value w i ;

[0033] The weighted centroid algorithm is used to determine the spatial center point (x). c ,y c ):

[0034]

[0035] Where n is the total number of selected pixels;

[0036] The three-dimensional spatial coordinates of the center point are mapped back to the spatial index of the two-dimensional slice in the three-dimensional volume data. Based on this spatial index, the slices closest to the spatial coordinates of the center point are selected in the sagittal and transverse directions, respectively.

[0037] Two-dimensional slices, sagittal slices, and transverse slices in the main direction are stitched together to obtain a multi-channel composite feature. This multi-channel composite feature contains image information of the same spatial location in three orthogonal directions, realizing the filtering of multi-layer spatial information and the fusion of high-dimensional features. This multi-channel composite feature will serve as the input for the subsequent multi-dimensional refinement module, providing rich spatial contextual information for refining the segmentation results.

[0038] Based on the preliminary segmentation probability map, information-rich slices are selected from different directions for further processing by the multi-dimensional refinement module. This not only achieves efficient integration of multi-directional and multi-modal information but also ensures smooth gradient flow, improving the interpretability and training stability of the model. Unlike existing simple feature splicing or full sequence processing, this method is more targeted and efficient.

[0039] As a preferred embodiment, a dual-module medical image segmentation method based on multi-directional MRI images further includes:

[0040] Based on the final segmentation probability map, the segmented lesions are re-calculated according to the original image size, and then back-located to the original multi-directional MRI image data. These coordinate points are marked, and their contrast or brightness is improved. Finally, the processed original multi-directional MRI image data are integrated together to obtain a 3D image model.

[0041] As a preferred option:

[0042] The preliminary segmentation and localization module includes a first encoder and a first decoder;

[0043] The two-dimensional slice in the main direction extracts multi-scale spatial features through the first encoder. The first decoder fuses the corresponding layer features of the first encoder through skip connections and gradually restores the spatial resolution through upsampling operations. After activation processing, the probability of the pixel belonging to the lesion area is displayed to obtain a preliminary segmentation probability map.

[0044] The multidimensional refinement module includes a second encoder and a second decoder;

[0045] The multi-channel composite feature input extracts high-dimensional features through the second encoder, and the second decoder fuses the corresponding layer features of the second encoder through skip connections, and gradually restores the spatial resolution through upsampling operations to complete the fine segmentation of the lesion.

[0046] The first encoder uses a pre-trained EfficientNet-B1 as its backbone network;

[0047] The first encoder consists of multiple convolutional blocks, each containing a convolutional layer, a batch normalization layer, and an activation function. Through convolution or pooling operations, the spatial resolution of the feature map is gradually reduced and the number of feature channels is increased.

[0048] The first decoder consists of multiple upsampling blocks. Each upsampling block first doubles the spatial size of the feature map through transposed convolution or upsampling operations, and then concatenates the feature map with the corresponding layer in the first encoder. The concatenated feature map then undergoes convolution, batch normalization and activation processing to gradually restore the spatial resolution and fuse low-level detail information with high-level semantic information.

[0049] To achieve the above objectives, the present invention also provides a dual-module medical image segmentation system based on multi-directional MRI images, comprising:

[0050] Preprocessing module: used to preprocess raw multi-directional MRI image data to obtain two-dimensional slices in the main direction;

[0051] Preliminary segmentation and localization module: Taking a two-dimensional slice in the main direction as input, it uses a deep learning network to perform preliminary segmentation and localization of the lesion and outputs a preliminary segmentation probability map;

[0052] Multi-layer information filtering and high-dimensional feature fusion module: Based on the preliminary segmentation probability map, obtain a slice in each of the other directions that is highly correlated with the spatial position of the segmentation region in the main direction, and stitch it with the two-dimensional slice in the main direction to obtain multi-channel composite features;

[0053] Multi-dimensional refinement module: The multi-channel composite feature input multi-dimensional refinement module performs multi-modal and multi-directional feature fusion and fine segmentation;

[0054] Loss function optimization module: The Dice loss function introduces outlier penalty, which applies additional penalties to segmentation edges and outliers.

[0055] Compared with the prior art, the beneficial effects of the present invention are:

[0056] (1) This invention proposes a novel dual-module structure that can effectively separate the localization and fine segmentation tasks, reduce the consumption of computing resources and the interference of redundant information, and improve the accuracy and robustness of segmentation.

[0057] The dual-module structure includes a preliminary segmentation and localization module and a multidimensional refinement module. The preliminary segmentation and localization module performs coarse localization of the lesion, and on this basis, the multidimensional refinement module performs deep fusion and fine segmentation of multimodal and multi-directional MRI information, resulting in high segmentation accuracy and good model interpretability.

[0058] (2) The residual connection structure realizes the direct information transmission between the initial segmentation and localization module and the multi-dimensional refined segmentation module, enabling the multi-dimensional refined segmentation module to further refine and correct the segmentation results based on the initial localization of the initial segmentation and localization module. In addition to enhancing the feature consistency between the initial segmentation and localization module and the multi-dimensional refined module, it also provides a direct gradient path for backpropagation, effectively alleviating the gradient vanishing problem in deep networks and ensuring the stability and performance of model training.

[0059] (3) To reduce the false positive rate, the Dice loss function (OPDL) with outlier penalty is introduced to impose additional penalties on segmentation edges and outliers, which effectively improves the accuracy of segmentation boundaries and overall performance.

[0060] (4) The multi-channel composite feature contains image information of the same spatial location in three orthogonal directions, realizing the filtering of multi-layer spatial information and the fusion of high-dimensional features; the multi-channel composite feature will be used as the input of the subsequent multi-dimensional refinement module, providing rich spatial context information for the refinement of the segmentation results; in addition, based on the preliminary segmentation probability map, information-rich slices are selected from different directions for further processing by the multi-dimensional refinement module; not only does it realize the efficient integration of multi-directional and multi-modal information, but it also ensures the smooth transmission of gradient flow, improving the interpretability and training stability of the model; unlike the existing simple feature splicing or full sequence processing, it has higher targeting and efficiency.

[0061] (5) Lightweight design: The overall system structure is lightweight, which is easy to deploy and apply in actual clinical environments, while taking into account both segmentation accuracy and computational efficiency.

[0062] (6) This invention addresses the challenge of segmenting MRI images of anal fistulas by proposing a segmentation method that combines dual-module phased processing with an innovative loss function. Extensive experimental verification demonstrates that this invention exhibits significant advantages in segmentation accuracy, false positive suppression, and model generalization ability. Attached Figure Description

[0063] Figure 1 This is a schematic diagram of the framework of a dual-module medical image segmentation method based on multi-directional MRI images according to the present invention;

[0064] Figure 2 This is a schematic diagram illustrating the visual effect of the present invention;

[0065] Figure 3 This is a schematic diagram of the 3D modeling effect of the present invention. Detailed Implementation

[0066] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0067] Application scenarios

[0068] This invention relates to the automatic segmentation of MRI images of anal fistulas, applied to clinical auxiliary diagnosis and surgical planning. Anal fistula is a common disease in which an abnormal passage forms between the perianal skin and the anal canal. Due to its high similarity to normal tissue structures, doctors are prone to misinterpretation when interpreting MRI images, especially in primary hospitals where experienced radiologists are lacking, resulting in high misdiagnosis and recurrence rates. MRI provides multi-directional and multimodal imaging information in coronal, sagittal, and transverse planes, making it an important tool for comprehensive assessment and surgical planning of anal fistulas, but it requires a high level of expertise from physicians. With the development of deep learning technology, automated and intelligent segmentation systems have become a key means to improve diagnostic efficiency and accuracy.

[0069] System Framework

[0070] This invention proposes a dual-module medical image segmentation system based on multi-directional MRI images, the overall structure of which is as follows: Figure 1 As shown.

[0071] The system mainly includes the following core modules:

[0072] Preliminary segmentation and localization module: This module takes a grayscale MRI image in a single main direction (such as the coronal plane) as input, and uses a deep learning network (such as U-Net structure, with EfficientNet-B1 encoder) to perform preliminary segmentation and localization of anal fistula lesions, and outputs a preliminary segmentation mask.

[0073] Multi-layer information filtering and high-dimensional feature fusion module: Based on the preliminary segmentation results of the preliminary segmentation and localization module, slices related to the lesion are automatically filtered from the other two directions (such as sagittal and transverse views) and stitched with the main direction image to form a three-channel input.

[0074] Multidimensional Refinement Module (MDR): The multidimensional refinement module performs multimodal and multidirectional feature fusion and fine segmentation on the three-channel input, and inputs the preliminary segmentation and the preliminary segmentation results of the localization module as residuals to further improve the segmentation accuracy.

[0075] Loss function optimization module: To reduce the false positive rate, the system introduces the Dice loss function (OPDL) with outlier penalty, which applies additional penalties to segmentation edges and outliers, thereby improving the accuracy and robustness of segmentation.

[0076] Compared with existing technologies, the advantages of this invention are:

[0077] Phased processing: Unlike existing technologies that directly process the entire sequence or all modal MRI images uniformly, this invention adopts a two-stage process of "coarse localization-fine segmentation". First, the main direction image is used for rapid localization, and then multi-directional information is fused for fine segmentation, which significantly reduces the consumption of computing resources and the interference of redundant information.

[0078] Information filtering and residual connection: Relevant slices are automatically filtered through the preliminary segmentation results, avoiding the inefficiency and feature redundancy of full sequence processing. The residual connection mechanism enables efficient information flow between modules, improving model interpretability and training stability.

[0079] Innovative Loss Function: To address the high false positive rate in medical image segmentation, an outlier penalty Dice loss function was designed, which effectively improves the accuracy of segmentation boundaries and overall performance.

[0080] Lightweight design: The overall system structure is lightweight, making it easy to deploy and apply in real-world clinical environments, while balancing segmentation accuracy and computational efficiency.

[0081] Application value

[0082] This system can be widely used in hospitals at all levels for automatic segmentation of MRI images of anal fistulas, assisting doctors in rapid and accurate lesion identification and surgical planning. It is especially suitable for primary healthcare institutions that lack professional radiologists and has significant clinical application value.

[0083] This invention also proposes a dual-module medical image segmentation method based on multi-directional MRI images. By combining multi-directional and multi-modal information fusion with innovative loss function design, it achieves efficient and accurate automatic lesion segmentation.

[0084] Its core methodology is as follows:

[0085] 1. Data Input and Preprocessing

[0086] This invention addresses the task of segmenting MRI images of anal fistulas. First, it systematically inputs and preprocesses the raw multi-directional MRI image data. All raw multi-directional MRI images are stored in standard medical imaging formats (such as DICOM or NIfTI). The system first performs unified format parsing on the image files, extracting T1-weighted sequence image data from the coronal, sagittal, and transverse planes, and converting them into a standard three-dimensional tensor structure.

[0087] Based on the spatial distribution characteristics of anal fistula lesions and the clinical diagnostic process, the coronal plane was preferentially selected as the primary input direction because it can more clearly display the structural layers of the anal canal and surrounding adipose tissue, which helps the model to initially locate the lesion. The other two directions serve as supplements for subsequent multidimensional information fusion.

[0088] For each subject, the system automatically traverses the 3D MRI volume data, unfolds the slices according to the main direction, generates a 2D slice sequence, and records the spatial index of each slice in the 3D volume data to facilitate subsequent spatial alignment and multi-directional information fusion.

[0089] To eliminate resolution and size differences caused by different patients and different scanning batches, all slices undergo size normalization and resampling. Specifically, the system first resamples the 3D volume data using trilinear interpolation based on the original voxel resolution, unifying its spatial resolution to [1.1364, 1.1364, 2.5]. Then, all 2D slices are adjusted to 512×512 pixels to ensure consistency in model input. To further eliminate inconsistencies in grayscale distribution caused by equipment and batch differences, the system performs grayscale normalization on each slice, linearly scaling pixel grayscale values ​​to the [0, 1] range. For slices with extreme values, percentile truncation (e.g., 1%–99%) is used to remove abnormally high and low grayscale values, improving the stability of data distribution.

[0090] During the training phase, to improve the model's generalization ability, the system can perform data augmentation operations on the input slices, including random rotation, translation, scaling, horizontal or vertical flipping, Gaussian noise perturbation, and random adjustment of contrast and brightness. All augmentation operations are performed only on the training set, while the validation and test sets maintain their original distribution.

[0091] At the same time, the system synchronously loads the manually labeled lesion segmentation mask corresponding to each two-dimensional slice, and performs size normalization and spatial resampling on the mask data in accordance with the image to ensure that the label and the image are completely aligned in space.

[0092] Each input image in the main direction needs to be localized by the preliminary segmentation module. For images in other directions, preprocessing is done together, but they are not input into the preliminary segmentation module.

[0093] After preprocessing, all image and label data are batch-organized according to patient number, orientation, slice index, and other information, and cached in an efficient data format for efficient reading during subsequent model training and inference stages. Through the above multi-step data input and preprocessing process, this invention ensures a high degree of consistency in spatial resolution, grayscale distribution, and size specifications of MRI image data from different sources and batches, laying a solid foundation for efficient training and accurate inference of the subsequent segmentation model.

[0094] 2. Preliminary segmentation and positioning module

[0095] In the first stage of the overall segmentation process, the input single grayscale 2D MRI slice in the main direction (coronal view) is first fed into the preliminary segmentation and localization module. This module employs a deep convolutional neural network based on the U-Net architecture, specifically including an encoder, a decoder, and multiple skip connections. The input 1×512×512 single-channel grayscale image first enters the encoder part, which uses a pre-trained EfficientNet-B1 as the backbone network. The encoder consists of multiple convolutional blocks, each containing convolutional layers, batch normalization layers, and activation functions (such as ReLU), and gradually reduces the spatial resolution and increases the number of feature channels through convolution or pooling operations with a stride of 2. After each convolutional block, the spatial size of the feature map is halved, and the number of channels is doubled, progressively extracting multi-scale spatial features of the image.

[0096] At each layer output of the encoder, a skip connection is established to connect the corresponding layer of the decoder. The decoder consists of a series of upsampling blocks. Each upsampling block first doubles the spatial size of the feature map through transposed convolution or upsampling operations, and then concatenates the feature map with the corresponding layer of the encoder. The concatenated feature map is then processed by convolution, batch normalization, and activation functions to gradually restore the spatial resolution and fuse low-level detail information with high-level semantic information. The final layer of the decoder outputs a feature map of the same size as the input image, which is then compressed to 1 channel by a 1×1 convolutional layer to obtain a single-channel segmentation probability map. This probability map is processed by a sigmoid activation function to normalize the output value of each pixel to the [0,1] interval, representing the probability that the pixel belongs to the lesion region. The final output preliminary segmentation probability map serves as the preliminary segmentation result, providing a basis for subsequent selection of spatial center points and multi-directional information fusion.

[0097] In the forward inference process of the entire preliminary segmentation and localization module, the input image sequentially undergoes multi-layer feature extraction by the encoder, information fusion of skip connections, spatial reconstruction by the decoder, and probability map generation to form a complete preliminary segmentation process.

[0098] 3. Multi-layer information filtering and high-dimensional feature fusion

[0099] After obtaining the preliminary segmentation probability map output by the preliminary segmentation and localization module, the system enters the multi-layer information filtering and high-dimensional feature fusion stage. First, the system performs pixel-level analysis on the preliminary segmentation probability map output by the preliminary segmentation and localization module, extracting the set of all pixels with a foreground probability greater than a set threshold (e.g., 0.5). For each selected pixel, the system records its spatial coordinates (x, y, y) in the two-dimensional slice. i ,y i) and its corresponding probability value w i This probability value is used as the weight for subsequent calculations of the spatial center point.

[0100] Next, the system uses a weighted centroid algorithm to determine the spatial pivot point of the segmented region. Specifically, the system takes the spatial coordinates and probability values ​​of all selected pixels as input to calculate the weighted centroid (x... c ,y c The calculation formula is as follows:

[0101]

[0102] Where $n$ is the total number of selected pixels, and w i Let w be the probability weight of the i-th pixel. To avoid low-confidence pixels interfering with the center point, the system adjusts the probability value w before calculation. i Threshold truncation is performed, retaining only pixels above a set threshold for calculation. If the current slice is a lesion-free or healthy sample, the initial segmentation and localization module outputs a low overall probability, and the system automatically sets the center point to the image center (256, 256), skipping iterative updates.

[0103] After determining the center point, the system maps the 3D spatial coordinates of that point back to the spatial index of the original 3D volumetric data. Using this spatial index as a reference, the system selects slices closest to the center point's spatial coordinates in both the sagittal and transverse directions. Specifically, in the 3D MRI volumetric data, keeping the center point coordinates constant, the system finds two slices closest to each other along the two directions other than the principal direction, calculates the Euclidean distance between each slice and the center point, and selects the slice with the smallest distance as the target slice. In this way, the system obtains one slice in each of the sagittal and transverse directions that is highly correlated with the spatial position of the segmented region in the principal direction.

[0104] Subsequently, the system stitches together the original main-direction slice, sagittal target slice, and transverse target slice according to channel dimension, forming a 3×512×512 multi-channel composite input. This composite input contains image information of the same spatial location in three orthogonal directions, realizing the filtering of multi-layer spatial information and the fusion of high-dimensional features. This multi-channel input will serve as the input for the subsequent multi-dimensional refinement module, providing rich spatial contextual information for refining the segmentation results. The entire process of multi-layer information filtering and high-dimensional feature fusion ensures a high degree of spatial correspondence and information complementarity between images in different directions, laying the foundation for improving subsequent segmentation accuracy.

[0105] 4. Multi-dimensional refined segmentation module

[0106] After multi-layer information filtering and high-dimensional feature fusion, the system enters the multi-dimensional refinement module. The input to this module is a 3×512×512 multi-channel composite image. The first channel is the original slice in the main direction (e.g., coronal), the second channel is the sagittal slice corresponding to the spatial center point of the segmented region, and the third channel is the corresponding transverse slice. The three channels are strictly aligned spatially to ensure that image information from the same spatial location in different directions can be effectively fused.

[0107] The multi-dimensional refinement module employs a deep convolutional neural network based on the U-Net architecture as its backbone. First, a three-channel composite image is input to the encoder of the multi-dimensional refinement module. The encoder consists of multiple convolutional blocks, each containing a convolutional layer, a batch normalization layer, and an activation function. Through convolution or pooling operations with a stride of 2, the spatial resolution is progressively reduced and the number of feature channels is increased, extracting high-dimensional features fused from multiple directions. The outputs of each encoder layer are connected to the corresponding layers of the decoder via skip connections. The decoder gradually restores the spatial resolution through upsampling operations and fuses low-level features from the encoder to achieve refined segmentation details.

[0108] In the multidimensional refinement module, a residual connection mechanism is introduced. Specifically, the preliminary segmentation probability map output by the preliminary segmentation and localization module, after being activated by Sigmoid, is directly input as a residual term to the end of the multidimensional refinement module decoder, and summed pixel-by-pixel with the output of the main branch of the multidimensional refinement module. This residual connection can be represented as:

[0109] P MDR =f MDR (I 3ch )+P PSL

[0110] Where P MDR f is the final segmentation probability map output by the multidimensional refinement module. MDR (I 3ch P represents the processing result of the three-channel input by the main branch of the multi-dimensional refinement module. PSL This is the initial segmentation probability map output by the initial segmentation and localization module. This structure enables direct information transfer between the initial segmentation and localization module and the multi-dimensional refinement module, allowing the multi-dimensional refinement module to further refine and correct the segmentation results based on the initial localization by the initial segmentation and localization module.

[0111] Residual connections not only enhance feature consistency between the initial segmentation and localization module and the multidimensional refinement module, but also provide a direct gradient path for backpropagation, effectively mitigating the gradient vanishing problem in deep networks and ensuring the stability and performance of model training. Finally, the multidimensional refinement module outputs a segmentation probability map of the same size as the input image, serving as the final segmentation result.

[0112] 5. Loss Function Optimization

[0113] During model training, we first divide each input image into blocks based on spatial priors. Taking a 512×512 image as an example, we divide it at 1 / 3 and 2 / 3 of its height and width, respectively, resulting in nine equal-sized 3×3 blocks. Then, for each pixel (i,j), we assign a weight w based on the block it belongs to. ij Edge block weight w edge The weight of the central block w center Satisfying w edge >w center The weight matrix w is generated once before training begins and remains unchanged throughout the training process.

[0114] During each forward propagation, the model outputs a predicted probability map Y = {y} ij}, with real labels Pixel-by-pixel correspondence. The calculation of the loss function involves the following steps:

[0115] 1. First, combine the predicted probability map $Y$ with the true labels. Multiplying element-wise with the weight matrix w yields the weighted prediction W⊙Y and the weighted label.

[0116] 2. Calculate the weighted intersection: That is, weights are only accumulated on pixels where both the prediction and label are positive.

[0117] 3. Calculate the sum of the weighted predicted region and the weighted true region respectively: $∑ i,j w ij y ij $ and

[0118] 4. Substitute the above results into the weighted Dice loss formula:

[0119]

[0120] During the loss backpropagation phase, the loss function predicts the value $y for each pixel. ij The gradient of $ is:

[0121]

[0122] In actual optimization, the pixels of the edge block are affected by w. ij The larger the pixel weight, the more its loss contribution and gradient are amplified. Thus, during training, the model prioritizes correcting misclassifications in these regions, especially false positives far from the lesion core. The central block has lower pixel weights, and the model prioritizes overall overlap and shape consistency in segmenting these regions.

[0123] The entire loss calculation process is performed independently for each image in each batch. The weight allocation scheme can be flexibly adjusted according to the spatial distribution characteristics of different datasets. For example, if the lesions are mainly distributed at the bottom of the image, the weight of the bottom block can be set to w. edge The remaining blocks are set to w. center The specific values ​​of the weights are treated as hyperparameters and are set in the experiment through cross-validation or empirical methods.

[0124] Through the aforementioned weighting mechanism, the loss function dynamically guides the model to focus on spatially more likely false positive regions with each parameter update, significantly improving the model's ability to suppress outlier misclassification while maintaining the segmentation accuracy of the target region. This method is simple to implement, computationally efficient, and adaptable to different tasks.

[0125] 6. Output and Application

[0126] After the above steps, we obtain a segmented image. We can then recalculate the segmentation coordinates of the identified lesions based on the original image size, reverse-locate them to the original medical image file, mark these coordinate points, improve their contrast or brightness, and then integrate the processed medical images to obtain a 3D image model. For example... Figure 3 As shown, doctors can use this model to more intuitively locate the position and shape of the lesion.

[0127] This invention addresses the challenge of segmenting MRI images of anal fistulas by proposing a segmentation method that combines dual-module, staged processing with an innovative loss function. Extensive experimental verification demonstrates significant advantages in segmentation accuracy, false positive suppression, and model generalization ability, as detailed below:

[0128] 1. Segmentation accuracy is significantly improved.

[0129] The method of this invention is compared with mainstream segmentation models such as Unet, Unet++, Manet, FPN, PSPNet, and LinkNet under five-fold cross-validation. The evaluation metrics are Dice score and IoU score. The experimental results are shown in Table 1, and the visualization effects of some mainstream segmentation models and this method are shown in the figure. Figure 2 :

[0130] Table 1: Dice score and IoU score of the present invention compared with mainstream segmentation methods

[0131] Method Dice (20 rounds) Dice (40 rounds) IoU (20 rounds) IoU (40 rounds) Unet 0.4943 0.5531 0.3787 0.4711 Unet++ 0.4967 0.6177 0.3792 0.5107 Manet 0.4883 0.6261 0.3759 0.5281 LinkNet 0.5038 0.5152 0.3868 0.4330 FPN 0.5037 0.5977 0.3888 0.4919 PSPNet 0.4691 0.6331 0.3507 0.5152 Invention 0.5517 0.7324 0.4302 0.5943

[0132] After 40 rounds of training, the method of this invention achieved a Dice score of 0.7324 and an IoU score of 0.5943, both significantly higher than the comparison model, fully demonstrating the effectiveness and advancement of this invention in the anal fistula segmentation task.

[0133] 2. False positives were significantly reduced, and the segmentation results were more accurate.

[0134] This invention innovatively introduces the outlier penalty Dice loss function (OPDL), which effectively suppresses false positives far from the lesion area. Quantitative analysis was performed using indicators such as Local False Positive (LFP) and Number of Segmentation Clusters (NC). As shown in Table 2, under the Mixed model, the LFP of OPDL was 0.0712, which is about 4% lower than the 0.1170 of the traditional Dice loss. The number of segmentation clusters was 1.1191, closer to the ideal value of 1, indicating that the segmentation results are more concentrated and accurate.

[0135] Table 2: Comparison of indices between OPDL and traditional Dice loss

[0136] Loss function Dice IOU LFP NC Dice 0.6117 0.4921 0.1170 1.0215 OPDL 0.7324 0.5943 0.0712 1.1191

[0137] 3. The dual-module structure improves the model's generalization ability and efficiency.

[0138] Ablation experiments verified the performance improvement of the dual-module structure (preliminary segmentation and localization module + multi-dimensional refinement module). As shown in Table 3, the Dice of the individual preliminary segmentation and localization module is 0.6271, and that of the multi-dimensional refinement module is 0.6635. After the dual-module integration, the Dice increased to 0.7324, representing an improvement of 16.8%. This structure not only improves segmentation accuracy but also effectively reduces the number of model parameters and computational resource consumption, facilitating practical deployment.

[0139] Table 3: Comparison of indicators under different numbers of modules

[0140]

[0141] 4. Excellent visualization results with outstanding boundary details.

[0142] By comparing the results of the visualization segmentation, the method of the present invention can accurately delineate the lesion area in the coronal, sagittal and transverse directions, with fewer false positives and clearer boundary details. It is particularly outstanding in small lesions and marginal areas, and has higher clinical practical value.

[0143] 5. Highly adaptable and easy to promote

[0144] The method of this invention has a lightweight structure, and the loss function can be flexibly adjusted according to different anatomical sites and data distributions. It has good generalization ability and prospects for widespread application, and is suitable for a variety of medical image segmentation scenarios.

[0145] While the invention has been described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the invention. Therefore, it should be understood that many modifications can be made to the exemplary embodiments, and other arrangements can be designed without departing from the spirit and scope of the invention as defined by the appended claims. It should be understood that different dependent claims and features herein can be combined in ways different from those described in the original claims. It is also understood that features described in conjunction with individual embodiments can be used in other embodiments.

Claims

1. A dual-module medical image segmentation method based on multi-directional MRI images, characterized in that, include: An image segmentation model is constructed, which includes a preliminary segmentation and localization module and a multi-dimensional refinement module; The raw multi-directional MRI image data is preprocessed, and the main direction is determined and two-dimensional slices of the main direction are obtained according to the task requirements. The two-dimensional slices in the main direction are input into the preliminary segmentation and localization module to perform preliminary segmentation and localization of the lesions, and output a preliminary segmentation probability map; Based on the preliminary segmentation probability map, obtain a slice in each of the other directions that is highly correlated with the spatial position of the segmented region in the main direction, and stitch it with the two-dimensional slice in the main direction to obtain a multi-channel composite feature. The multi-channel composite features are input into the multi-dimensional refinement module for deep fusion and fine segmentation, thereby completing the fine segmentation of the lesions.

2. The dual-module medical image segmentation method based on multi-directional MRI images according to claim 1, characterized in that, Also includes: After being activated by Sigmoid, the preliminary segmentation probability map is input as the residual term to the end of the multidimensional refinement module. It is then summed pixel by pixel with the output of the main branch of the multidimensional refinement module to obtain the final segmentation probability map.

3. The dual-module medical image segmentation method based on multi-directional MRI images according to claim 1, characterized in that, Also includes: The model is trained using the Dice loss function, which introduces outlier penalties, to penalize segmentation edges and outliers.

4. The dual-module medical image segmentation method based on multi-directional MRI images according to claim 1, characterized in that, The preprocessing of the raw multi-directional MRI image data, determining the principal direction according to task requirements, and obtaining two-dimensional slices of the principal direction includes the following steps: The original multi-directional MRI image data was parsed in a unified format to extract the coronal, sagittal and transverse image data, and then converted into standard three-dimensional volume data. Based on the spatial distribution characteristics of anal fistula lesions and the clinical diagnostic process, the coronal plane was selected as the main direction. Traverse the 3D volume data, slice and unfold according to the main direction to generate a 2D slice sequence, and record the spatial index of each 2D slice in the 3D volume data.

5. The dual-module medical image segmentation method based on multi-directional MRI images according to claim 4, characterized in that, The preprocessing of the raw multi-directional MRI image data, including determining the principal direction and obtaining two-dimensional slices of the principal direction according to task requirements, further includes the following steps: All 2D slices are subjected to size normalization and resampling. The grayscale of the two-dimensional slices after size normalization and resampling is normalized. The corresponding labels are obtained by synchronously loading the manually annotated lesion segmentation mask corresponding to each two-dimensional slice.

6. The dual-module medical image segmentation method based on multi-directional MRI images according to claim 4, characterized in that, The step of obtaining a slice in each of the other directions based on the preliminary segmentation probability map, which is highly correlated with the spatial position of the segmented region in the main direction, and then stitching it with the two-dimensional slice in the main direction to obtain a multi-channel composite feature includes the following steps: Perform pixel-level analysis on the preliminary segmentation probability map and select the set of all pixels whose foreground probability is greater than a set threshold; For each selected pixel, record its spatial coordinates (x, y) in the corresponding 2D slice. i ,y i ) and its corresponding probability value w i ; The weighted centroid algorithm is used to determine the spatial center point (x). c ,y c ): Where n is the total number of selected pixels; The three-dimensional spatial coordinates of the center point are mapped back to the spatial index of the two-dimensional slice in the three-dimensional volume data. Based on this spatial index, the slices closest to the spatial coordinates of the center point are selected in the sagittal and transverse directions, respectively. By stitching together the two-dimensional slices, sagittal slices, and transverse slices along the main direction, a multi-channel composite feature is obtained.

7. The dual-module medical image segmentation method based on multi-directional MRI images according to claim 2, characterized in that, Also includes: Based on the final segmentation probability map, the segmented lesions are re-calculated according to the original image size, and then back-located to the original multi-directional MRI image data. These coordinate points are marked, and their contrast or brightness is improved. Finally, the processed original multi-directional MRI image data are integrated together to obtain a 3D image model.

8. The dual-module medical image segmentation method based on multi-directional MRI images according to claim 1, characterized in that: The preliminary segmentation and localization module includes a first encoder and a first decoder; The two-dimensional slice in the main direction extracts multi-scale spatial features through the first encoder. The first decoder fuses the corresponding layer features of the first encoder through skip connections and gradually restores the spatial resolution through upsampling operations. After activation processing, the probability of the pixel belonging to the lesion area is displayed to obtain a preliminary segmentation probability map. The multidimensional refinement module includes a second encoder and a second decoder; The multi-channel composite feature input extracts high-dimensional features through the second encoder, and the second decoder fuses the corresponding layer features of the second encoder through skip connections, and gradually restores the spatial resolution through upsampling operations to complete the fine segmentation of the lesion.

9. A dual-module medical image segmentation method based on multi-directional MRI images according to claim 8, characterized in that: The first encoder uses a pre-trained EfficientNet-B1 as its backbone network; The first encoder consists of multiple convolutional blocks, each containing a convolutional layer, a batch normalization layer, and an activation function. Through convolution or pooling operations, the spatial resolution of the feature map is gradually reduced and the number of feature channels is increased. The first decoder consists of multiple upsampling blocks. Each upsampling block first doubles the spatial size of the feature map through transposed convolution or upsampling operations, and then concatenates the feature map with the corresponding layer in the first encoder. The concatenated feature map then undergoes convolution, batch normalization and activation processing to gradually restore the spatial resolution and fuse low-level detail information with high-level semantic information.

10. A dual-module medical image segmentation system based on multi-directional MRI images, characterized in that, include: Preprocessing module: used to preprocess raw multi-directional MRI image data to obtain two-dimensional slices in the main direction; Preliminary segmentation and localization module: Taking a two-dimensional slice in the main direction as input, it uses a deep learning network to perform preliminary segmentation and localization of the lesion and outputs a preliminary segmentation probability map; Multi-layer information filtering and high-dimensional feature fusion module: Based on the preliminary segmentation probability map, obtain a slice in each of the other directions that is highly correlated with the spatial position of the segmentation region in the main direction, and stitch it with the two-dimensional slice in the main direction to obtain multi-channel composite features; Multi-dimensional refinement module: The multi-channel composite feature input multi-dimensional refinement module performs multi-modal and multi-directional feature fusion and fine segmentation; Loss function optimization module: The Dice loss function introduces outlier penalty, which applies additional penalties to segmentation edges and outliers.