Douglas fossa occlusion prediction method based on multi-modal MRI (Magnetic Resonance Imaging) feature fusion

By applying multimodal MRI feature fusion and 3D-Mamba network, the problem of insufficient accuracy in predicting POD occlusion in existing models has been solved, enabling early, non-invasive, and objective diagnosis of POD occlusion, thus improving diagnostic efficiency and accuracy.

CN121962718APending Publication Date: 2026-05-01THE FIRST PEOPLES HOSPITAL OF XIAOSHAN DISTRICT HANGZHOU +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
THE FIRST PEOPLES HOSPITAL OF XIAOSHAN DISTRICT HANGZHOU
Filing Date
2025-12-31
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing AI-based POD occlusion prediction models do not make sufficient use of T1 and T2 weighted MRI image features, neglect the continuity of lesions in 3D space and overall morphological changes, resulting in insufficient diagnostic accuracy and reliance on physician experience, lacking objectivity.

Method used

A multimodal MRI feature fusion method was adopted, which integrates 3DT1 and T2 weighted MRI images, extracts features using a 3D-Mamba network, and designs a dual attention mechanism for feature fusion to achieve adaptive feature fusion of anatomy and pathophysiology, thereby improving the accuracy of POD occlusion prediction.

Benefits of technology

It enables early, non-invasive diagnosis of POD occlusion, reduces reliance on invasive laparoscopic examinations, improves the objectivity, accuracy, and efficiency of diagnosis, and reduces the risk of diagnostic delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121962718A_ABST
    Figure CN121962718A_ABST
Patent Text Reader

Abstract

The invention provides a Douglas fossa occlusion prediction method based on multi-mode MRI (Magnetic Resonance Imaging) feature fusion. The method comprises the following steps: collecting multi-modal 3D MRI data of a patient, wherein the multi-modal 3D MRI data comprises pelvic cavity 3D T1 weighted and T2 weighted MRI image data; the method comprises the following steps of: firstly, carrying out multi-modal MRI (Magnetic Resonance Imaging), preprocessing the multi-modal MRI to realize image registration, mapping the multi-modal MRI to the same space, and secondly, realizing local contrast enhancement and marginal definition optimization of the MRI image through adaptive histogram equalization; secondly, designing a parallel-based double-branch 3D-Mama network to extract long-term spatial dependence and continuous features of the weighted MRI images T1 and T2; a 3D channel and space double attention fusion module is further designed, and multi-mode MRI feature deep fusion and adaptive feature selection are achieved; and finally, mapping the learned multi-modal MRI fusion features to a classification space by using a multi-layer perceptron, and carrying out POD occlusion classification prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital medical technology, and more specifically to a method for predicting Douglas's fossa occlusion based on multimodal MRI feature fusion. Background Technology

[0002] Endometriosis (EM) is characterized by the abnormal growth of endometrial-like tissue outside the uterine cavity, typically causing symptoms such as chronic pain, prolonged menstrual periods, and infertility. The incidence of this disease is as high as 10% in women of reproductive age. EM presents with diverse and nonspecific clinical manifestations, and current diagnosis still heavily relies on invasive procedures such as laparoscopy, which can easily lead to potential complications such as postoperative adhesions and infections. Furthermore, EM diagnosis is time-consuming; on average, it takes 6.4 years from the onset of symptoms to a definitive diagnosis. This significant diagnostic delay not only increases the risk of disease progression but also misses the optimal intervention window to protect fertility, impacting the patient's quality of life.

[0003] The pouch of Douglas (POD) is the lowest point of the female peritoneum between the rectum and uterus. It is a site where pelvic effusion, inflammatory cells, and exfoliated endometrial cells accumulate, making it one of the most commonly affected areas in endometrial lesions (EM). POD occlusion is one of the most important characteristics of EM. T1-weighted and T2-weighted MRI sequences can clearly present various signs of EM and are key evidence for determining whether the POD is occluded. However, current assessment of POD occlusion is highly dependent on physician experience and is subjective. Therefore, by integrating patients' T1- and T2-weighted MRI images, an artificial intelligence model capable of automatically identifying PODs can be developed to improve the accuracy of POD occlusion diagnosis and achieve early diagnosis of EM. Current AI-based POD occlusion prediction models do not fully utilize the features of T1- and T2-weighted MRI images, mostly using only one type for analysis, which limits the predictive power of a single MRI aspect. Furthermore, existing models mostly use MRI slices for feature extraction, neglecting the continuity of the lesion in 3D space and overall morphological changes, thus affecting the model's effectiveness. Summary of the Invention

[0004] To address the shortcomings of existing technologies, the present invention aims to provide a method for predicting Douglas pit occlusion based on multimodal MRI feature fusion. This method integrates 3D T1 and T2 weighted MRI images, extracts long-range dependent features of T1 and T2 weighted MRI using 3D-Mamba, and further designs a dual attention mechanism to achieve deep fusion of multimodal MRI features, enabling adaptive feature fusion of anatomy and pathophysiology, thereby improving the accuracy of POD occlusion prediction.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for predicting Douglas's fossa occlusion based on multimodal MRI feature fusion, comprising the following steps: Step 1: Collect multimodal 3D MRI data, including 3D T1-weighted and T2-weighted MRI images of the patient's pelvis; Step 2: Preprocess the medical images collected in Step 1, performing multimodal MRI image registration and local contrast enhancement; Step 3: Extract T1- and T2-weighted MRI image features based on a parallel dual-branch 3D-Mamba network; Step 4: Construct a 3D channel and spatial dual attention fusion module, and then use the constructed 3D channel and spatial dual attention fusion module to achieve deep fusion of multimodal MRI features; Step 5: Use a multilayer perceptron to map the learned multimodal MRI fusion features to the classification space to perform POD occlusion classification prediction.

[0006] As a further improvement of the present invention, the pelvic 3DT1-weighted and T2-weighted MRI image data of the patient collected in step one of the Douglas fossa occlusion prediction method based on multimodal MRI feature fusion are as follows: T1-weighted imaging: It can clearly show the normal anatomical structure of pelvic organs and is the main sequence for determining whether the anterior wall of the rectum is stretched or adhered. T2-weighted imaging primarily reflects the water content of tissues, excels at displaying fluids and soft tissues, and is extremely sensitive to changes in water content, which is helpful in observing tissue lesions.

[0007] As a further improvement of the present invention, the multimodal MRI image registration in step two of the Douglas fossa occlusion prediction method based on multimodal MRI feature fusion is as follows: using T2-weighted imaging as the reference image, T1-weighted imaging is registered to the same space: in, It is a mutual information similarity measurement function, after registration. for .

[0008] As a further improvement of the present invention, the specific method for extracting T1 and T2 weighted MRI image features in step three of the Douglas fossa occlusion prediction method based on multimodal MRI feature fusion is as follows: A 3D convolutional neural network is used to extract the spatial continuity information of the lesion, and the corresponding 3D-patch features are obtained through linear mapping: in, Represents 3D-CNN operations. The weight matrix is ​​a learnable matrix. For the corresponding bias parameters, The final learned 3D-patch features; Add learnable position vectors to the extracted 3D-patch features The feature vector sequence obtained from the input 3D-Mamba network is: ; After obtaining the block embedding vector sequence of three-dimensional medical images, a 3D-Mamba network based on a state-space model is used to model the long-range spatial-sequence dependencies within the feature sequence.

[0009] As a further improvement of the present invention, the 3D-Mamba network in step three of the Douglas fossa occlusion prediction method based on multimodal MRI feature fusion includes two branches. The first branch sequentially performs linear transformation, activation, and convolution enhancement on the input features, and realizes dynamic state propagation across sequence positions through a structured state-space model to obtain the final features. The second branch uses a linear mapping to generate a gated signal. , used to control the output of SSM; in, and These are the weight matrix and the bias vector, respectively; combining the first and second branches, the final feature obtained is: in, Represents the dot product operation, and linear(*) is the linear projection layer, which is learned. , where c is the number of channels.

[0010] As a further improvement of the present invention, the first branch of the Douglas fossa occlusion prediction method based on multimodal MRI feature fusion sequentially performs linear transformation, activation, and convolution enhancement on the input features, and realizes dynamic state propagation across sequence positions through a structured state-space model in the following specific way: First, process the input feature vector sequence Applying a linear mapping and the SiLU activation function yields intermediate feature representations: in, This is the weight matrix. Let SiLU be the bias vector, and let SiLU be the activation function, defined as follows: Subsequently, regarding Three-dimensional convolution operations are performed to enhance local context modeling capabilities. The convolution output is then fed into a 3D-SSM suitable for 3D images, effectively capturing long-range spatial dependencies in 3D-MRI block sequences and modeling the overall anatomical relationship between the Douglas fossa region and surrounding tissues. Then, modeling is performed in the order of forward space, backward space, forward sequence, and backward sequence. The 3D input features are flattened into four sequences, and the four sequences are fed into 3D-SSM respectively to learn the global dependency features of the 3D image from different perspectives. Subsequently, the features obtained from scanning in the four directions are weighted and fused to obtain the final features. .

[0011] As a further improvement of the present invention, in the Douglas fossa occlusion prediction method based on multimodal MRI feature fusion, after flattening the three-dimensional input features into four sequences, for a certain sequence, the hidden state of the previous time step is used. and current input Based on this, iteratively update the state and output: in, It is the input patch at position t in the input sequence. It is the hidden state at time t, which carries the historical information of the sequence; and To obtain the discretized system matrix using the zero-order preservation technique, This is the projection matrix, and its calculation method is as follows: in, This represents the discretization step size, which indicates the phased preservation of the input. It is an identity matrix.

[0012] As a further improvement of the present invention, the specific method for achieving deep fusion of multimodal MRI features in step four of the Douglas fossa occlusion prediction method based on multimodal MRI feature fusion is as follows: First, the feature maps of T1 and T2 are... and Features are obtained by concatenating along the channel dimension. : ; Then 3D channel attention is achieved through... Along spatial dimensions Global average pooling is performed to obtain channel description vectors that compress global spatial information. Next, the nonlinear interactions between channels are learned through two linear mapping layers, and channel weight vectors are output to measure the importance of different channels for POD occlusion prediction. in , r is the compression ratio, and the final result is... This refers to the importance weight of each channel; simultaneously, 3D spatial attention is applied along the channel dimension. Simultaneous average pooling and max pooling are performed to obtain two 3D feature maps. and Average pooling preserves overall background information, while max pooling captures the most salient features. These two feature maps are then concatenated along the channel dimension and processed using a convolutional kernel with a size of [missing information]. The 3D-CNN analyzes the stitched feature map to learn the relative importance of spatial positions, and then generates spatial attention weights through the sigmoid activation function. The specific calculation method is as follows: Finally, channel weights and spatial weights are applied simultaneously to the stitched feature map, resulting in the fusion features of the two modalities of MRI: .

[0013] As a further improvement of the present invention, the specific method for POD occlusion classification and prediction in step five of the Douglas fossa occlusion prediction method based on multimodal MRI feature fusion is as follows: After obtaining the fused features from multimodal MRI, they are flattened into a one-dimensional feature vector. This vector is then input into a multilayer perceptron (MLP) classifier to predict POD occlusion. This MLP typically consists of one or two fully connected layers, using the ReLU activation function in between, and finally outputting a probability value between 0 and 1 through a sigmoid function. The details are as follows: in, This is the probability of POD occlusion predicted by the model. The larger the value, the greater the probability that the POD is in an occluded state, and vice versa.

[0014] The beneficial effects of this invention are as follows: This invention utilizes the 3D-Mamba model to achieve long-distance spatial dependence modeling on 3D MRI data, realizing the extraction of continuous and holistic spatial features of 3D images, thereby providing a more comprehensive assessment of pelvic anatomical structures; furthermore, through a dual attention mechanism, it achieves deep fusion of multimodal MRI features based on medical prior knowledge, effectively utilizing the complementary information between T1 and T2 modalities, and ultimately providing an objective and accurate POD occlusion prediction model, reducing reliance on invasive laparoscopic examinations and achieving early non-invasive diagnosis. Attached Figure Description

[0015] Figure 1 This is a flowchart illustrating the Douglas fossa occlusion prediction method based on modal MRI feature fusion of the present invention. Figure 2 This is a schematic diagram of 3D-Mamba scanning in four directions in this invention. Detailed Implementation

[0016] The present invention will now be described in further detail with reference to the embodiments shown in the accompanying drawings.

[0017] Reference Figures 1 to 2 As shown, a method for predicting Douglas's fossa occlusion based on multimodal MRI feature fusion is described, and the method steps are as follows: Step 1: Multimodal 3D MRI data collection, acquiring 3D T1-weighted and T2-weighted MRI images of the patient's pelvis; T1 and T2 are contrast modes in MRI that reflect different tissue characteristics. By collecting MRI images in both modes, we can comprehensively reflect the anatomical and pathological information of the tissue, providing a basis for the diagnosis, localization and characterization of diseases.

[0018] T1-weighted imaging: It can clearly show the normal anatomical structure of pelvic organs and is the main sequence for determining whether the anterior wall of the rectum is stretched or adhered.

[0019] T2-weighted imaging primarily reflects the water content of tissues, excels at displaying fluids and soft tissues, and is extremely sensitive to changes in water content, which is helpful in observing tissue lesions.

[0020] Step 2: Medical image preprocessing, multimodal MRI image registration and local contrast enhancement; MRI preprocessing is an indispensable and crucial process connecting raw data with advanced analysis. Through MRI image preprocessing, image quality can be improved, and all data can be standardized into a unified spatiotemporal space to facilitate subsequent quantitative analysis, inter-group comparisons, and modeling analysis.

[0021] For MRI image registration, T2-weighted imaging is used as the reference image, and T1-weighted imaging is registered to the same space; the main goal is to find a spatial transformation. To align the two modalities and minimize their differences, specifically: in It is a mutual information similarity measurement function, after registration. for Image registration can correct spatial positional deviations between T1 and T2 images caused by minute patient movements during scanning, ensuring that T1 and T2 signals from the same anatomical location correspond precisely in space.

[0022] Furthermore, due to the complexity of pelvic tissues and the large dynamic range of signal intensity, adaptive histogram equalization is employed to enhance local contrast. Adaptive histogram equalization calculates and equalizes the histogram of the local neighborhood of each voxel, thereby enhancing local image contrast and effectively improving the edge clarity of fine structures such as the uterosacral ligament and rectovaginal septum, making the subtle texture features of early fibrotic adhesions more prominent.

[0023] Step 3: Extract T1 and T2 weighted MRI image features based on a parallel dual-branch 3D-Mamba network; POD occlusion is a large-scale, multi-tissue, three-dimensional spatial event. Therefore, by using a 3D-Mamba network to selectively focus on multiple key, geographically distant anatomical regions in the entire 3D image and establish corresponding spatial-pathological connections, we can learn the global and complex pathological patterns that lead to POD occlusion.

[0024] EM lesions are three-dimensional and may present different morphologies on slices at different depths. Before transmitting MRI data into the Mamba network, the 3D MRI data needs to be converted into a series of patch markers for the preprocessed 3D T1 MRI sequence data. In other words, its segmented 3D-patch can be represented as Where N is the total number of 3D-patch. Considering that treating 3D MRI data as a series of continuous 2D slices would destroy its inherent spatial continuity, a 3D convolutional neural network (3D-CNN) is used here to extract the spatial continuity information of the lesions, and the corresponding 3D-patch features are obtained through linear mapping. Specifically: in, Represents 3D-CNN operations. The weight matrix is ​​a learnable matrix. For the corresponding bias parameters, The final learned 3D-patch features retain spatial context information from the 3D MRI volume data in the coronal, sagittal, and axial dimensions, which helps in understanding the spatial orientation and infiltration depth of adherent tissues. Furthermore, to maintain the positional relationships between different 3D-patches, learnable position vectors are added to them. Finally, the feature vector sequence input into the 3D-Mamba network is .

[0025] After obtaining the block embedding vector sequence of 3D medical images, a 3D-Mamba network based on a state-space model is used to model the long-range spatial-sequence dependencies within the feature sequence. This network mainly consists of two branches. The first branch sequentially performs linear transformations, activations, and convolutional enhancements on the input features, and uses a structured state-space model (SSM) to achieve dynamic state propagation across sequence positions. Specifically, the input feature vector sequence is first processed... Applying a linear mapping and the SiLU activation function yields intermediate feature representations: in, This is the weight matrix. Let SiLU be the bias vector, and let SiLU be the activation function, defined as follows: Subsequently, regarding Three-dimensional convolution operations are performed to enhance local context modeling capabilities. The convolution output is then fed into a 3D-SSM suitable for 3D images, effectively capturing long-range spatial dependencies in 3D-MRI block sequences and modeling the overall anatomical relationship between the Douglas fossa region and surrounding tissues. To enhance the global information interaction capabilities of 3D-Mamba, modeling is performed in the order of forward space, backward space, forward sequence, and backward sequence to help it adapt to 3D image information. Through the aforementioned process, the 3D input features can be flattened into four sequences, which are then fed into 3D-SSM to learn the global dependency features of the 3D image from different perspectives. For any given sequence, the hidden state from the previous time step is used. and current input Based on this, iteratively update the state and output: in It is the input patch at position t in the input sequence. It is the hidden state at time t, which carries the historical information of the sequence. and To obtain the discretized system matrix using the zero-order preservation technique, This is the projection matrix, and its calculation method is as follows: in This represents the discretization step size, which indicates the phased preservation of the input. It is the identity matrix; unlike traditional sequence-state models that use static parameterization, it allows modification of the projection matrix based on the input, thereby enabling selective attention to each sequence unit. Specifically, parameterized B, C, and It can be used to achieve selective attention to the state of a sequence.

[0026] Subsequently, the features obtained from scanning in four directions are weighted and fused to obtain the final features. This is used for subsequent feature analysis.

[0027] The second branch of 3D-Mamba uses linear mapping to generate gating signals. , used to control the output of SSM; in, and These are the weight matrix and the bias vector, respectively, and the final features obtained by 3D-Mamba are: in Represents the dot product operation, and linear(*) is the linear projection layer, which is learned. Where c is the number of channels. The 3D-Mamba network enables the model to dynamically focus on key features during image scanning and understand the long-range correlations of these features in three-dimensional space, thereby achieving context-aware feature encoding of POD occlusion states. Similarly, for 3D T2-weighted MRI images, the corresponding features can be obtained through the 3D-Mamba network. .

[0028] Step 4: Design a 3D channel and spatial dual attention fusion module to achieve deep fusion of multimodal MRI features; After obtaining the 3D T1 and T2 weighted MRI image features, the T1 and T2 feature maps are... and Features are obtained by concatenating along the channel dimension. : 3D channel attention through Along spatial dimensions Global average pooling (GAP) is performed to obtain channel description vectors that compress global spatial information. Furthermore, the nonlinear interactions between channels are learned through two linear mapping layers, and channel weight vectors are output to measure the importance of different channels for POD occlusion prediction. Specifically: in , r is the compression ratio, and the final result is... That is, the importance weight of each channel.

[0029] 3D spatial attention along the channel dimension Simultaneous average pooling and max pooling are performed to obtain two 3D feature maps. and Average pooling preserves overall background information, while max pooling captures the most salient features. These two feature maps are then concatenated along the channel dimension and processed using a convolutional kernel with a size of [missing information]. The 3D-CNN analyzes the stitched feature map to learn the relative importance of spatial positions, and then generates spatial attention weights through the sigmoid activation function. The specific calculation method is as follows: The 3D spatial attention mechanism can be used to locate the three-dimensional spatial region most relevant to POD occlusion. Finally, channel weights and spatial weights are simultaneously applied to the stitched feature map to achieve fine-grained, adaptive feature selection, ultimately resulting in the fused MRI features of the two modalities: Step 5: Use a multilayer perceptron to map the learned multimodal MRI fusion features to the classification space to perform POD occlusion classification prediction.

[0030] After obtaining the fused features from multimodal MRI, these features are flattened into a one-dimensional feature vector. This vector is then input into a multilayer perceptron (MLP) classifier to predict POD occlusion. The MLP typically consists of one or two fully connected layers, using the ReLU activation function in between, and finally outputting a probability value between 0 and 1 via a sigmoid function. Specifically: in This is the probability of POD occlusion predicted by the model. The larger the value, the greater the probability that the POD is in an occluded state, and vice versa.

[0031] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A method for predicting Douglas's fossa occlusion based on multimodal MRI feature fusion, characterized in that: Includes the following steps: Step 1: Collect multimodal 3D MRI data, including acquiring 3D T1-weighted and T2-weighted MRI images of the patient's pelvis; Step 2: Preprocess the medical images collected in Step 1, performing multimodal MRI image registration and local contrast enhancement; Step 3: Extract T1 and T2 weighted MRI image features based on a parallel dual-branch 3D-Mamba network; Step 4: Construct a 3D channel and spatial dual attention fusion module, and then use the constructed 3D channel and spatial dual attention fusion module to achieve deep fusion of multimodal MRI features; Step 5: Use a multilayer perceptron to map the learned multimodal MRI fusion features to the classification space to perform POD occlusion classification prediction.

2. The Douglas fossa occlusion prediction method based on multimodal MRI feature fusion according to claim 1, characterized in that: The specific 3D T1-weighted and T2-weighted MRI images of the patient's pelvis collected in step one are as follows: T1-weighted imaging: It can clearly show the normal anatomical structure of pelvic organs and is the main sequence for determining whether the anterior wall of the rectum is stretched or adhered. T2-weighted imaging primarily reflects the water content of tissues, excels at displaying fluids and soft tissues, and is extremely sensitive to changes in water content, which is helpful in observing tissue lesions.

3. The Douglas fossa occlusion prediction method based on multimodal MRI feature fusion according to claim 1 or 2, characterized in that: The multimodal MRI image registration in step two is as follows: using T2-weighted imaging as the reference image, the T1-weighted imaging is registered to the same space: in, It is a mutual information similarity measurement function, after registration. for .

4. The Douglas fossa occlusion prediction method based on multimodal MRI feature fusion according to claim 1 or 2, characterized in that: The specific method for extracting T1 and T2 weighted MRI image features in step three is as follows: A 3D convolutional neural network is used to extract the spatial continuity information of the lesion, and the corresponding 3D-patch features are obtained through linear mapping. in, Represents 3D-CNN operations. The weight matrix is ​​a learnable matrix. For the corresponding bias parameters, The final learned 3D-patch features; Add learnable position vectors to the extracted 3D-patch features The feature vector sequence obtained from the input 3D-Mamba network is: ; After obtaining the block embedding vector sequence of three-dimensional medical images, a 3D-Mamba network based on a state-space model is used to model the long-range spatial-sequence dependencies within the feature sequence.

5. The Douglas fossa occlusion prediction method based on multimodal MRI feature fusion according to claim 4, characterized in that: The 3D-Mamba network in step three includes two branches. The first branch sequentially performs linear transformation, activation, and convolution enhancement on the input features, and uses a structured state-space model to achieve dynamic state propagation across sequence positions to obtain the final features. The second branch uses a linear mapping to generate a gated signal. , used to control the output of SSM; in, and These are the weight matrix and the bias vector, respectively; combining the first and second branches, the final feature obtained is: in, Represents the dot product operation, and linear(*) is the linear projection layer, which is learned. , where c is the number of channels.

6. The Douglas fossa occlusion prediction method based on multimodal MRI feature fusion according to claim 5, characterized in that: The first branch sequentially performs linear transformation, activation, and convolution enhancement on the input features, and achieves dynamic state propagation across sequence positions through a structured state-space model in the following specific way: First, process the input feature vector sequence Applying a linear mapping and the SiLU activation function yields intermediate feature representations: in, This is the weight matrix. Let SiLU be the bias vector, and let SiLU be the activation function, defined as follows: Subsequently, regarding Three-dimensional convolution operations are performed to enhance local context modeling capabilities. The convolution output is then fed into a 3D-SSM suitable for 3D images, effectively capturing long-range spatial dependencies in 3D-MRI block sequences and modeling the overall anatomical relationship between the Douglas fossa region and surrounding tissues. Then, modeling is performed in the order of forward space, backward space, forward sequence, and backward sequence. The 3D input features are flattened into four sequences, and the four sequences are fed into 3D-SSM respectively to learn the global dependency features of the 3D image from different perspectives. Subsequently, the features obtained from scanning in the four directions are weighted and fused to obtain the final features. .

7. The Douglas fossa occlusion prediction method based on multimodal MRI feature fusion according to claim 6, characterized in that: After flattening the 3D input features into four sequences, for any one of these sequences, the hidden state from the previous time step is used. and current input Based on this, iteratively update the state and output: in, It is the input patch at position t in the input sequence. It is the hidden state at time t, which carries the historical information of the sequence; and To obtain the discretized system matrix using the zero-order preservation technique, This is the projection matrix, and its calculation method is as follows: in, This represents the discretization step size, which indicates the phased preservation of the input. It is an identity matrix.

8. The Douglas fossa occlusion prediction method based on multimodal MRI feature fusion according to claim 1 or 2, characterized in that: The specific method for achieving multimodal MRI feature deep fusion using the constructed 3D channel and spatial dual attention fusion module in step four is as follows: First, the feature maps of T1 and T2 are... and Features are obtained by concatenating along the channel dimension. : ; Then 3D channel attention is achieved through... Along spatial dimensions Global average pooling is performed to obtain channel description vectors that compress global spatial information. Next, the nonlinear interactions between channels are learned through two linear mapping layers, and channel weight vectors are output to measure the importance of different channels for POD occlusion prediction. in , r is the compression ratio, and the final result is... This refers to the importance weight of each channel; simultaneously, 3D spatial attention is applied along the channel dimension. Simultaneous average pooling and max pooling are performed to obtain two 3D feature maps. and Average pooling preserves overall background information, while max pooling captures the most salient features; Then, these two feature maps are concatenated along the channel dimension and a convolution kernel of size [size missing] is applied. The 3D-CNN analyzes the stitched feature map to learn the relative importance of spatial positions, and then generates spatial attention weights through the sigmoid activation function. The specific calculation method is as follows: Finally, channel weights and spatial weights are applied simultaneously to the stitched feature map, resulting in the fusion features of the two modalities of MRI: 。 9. The Douglas fossa occlusion prediction method based on multimodal MRI feature fusion according to claim 1 or 2, characterized in that: The specific method for POD occlusion classification and prediction in step five is as follows: After obtaining the fused features from multimodal MRI, they are flattened into a one-dimensional feature vector. This vector is then input into a multilayer perceptron (MLP) classifier to predict POD occlusion. This MLP typically consists of one or two fully connected layers, using the ReLU activation function in between, and finally outputting a probability value between 0 and 1 through a sigmoid function. The details are as follows: in, This is the probability of POD occlusion predicted by the model. The larger the value, the greater the probability that the POD is in an occluded state, and vice versa.