Multi-module fusion segmentation network implementation method for breast cancer DCE-MRI image segmentation

By introducing a multi-module fusion segmentation network - a harmonic visual python framework in medical image segmentation, combining frequency domain and spatial domain feature extraction, attention mechanism and multi-scale convolution, the limitations of the existing technology in complex backgrounds and small lesions are solved, and high-precision and robust breast cancer image segmentation are achieved.

CN119963545AInactive Publication Date: 2025-05-09HANGZHOU DIANZI UNIV

Patent Information

Application Number
CN202510429467.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing medical imaging segmentation methods have limitations in dealing with complex backgrounds, noise interference and edge feature ambiguity, especially in the segmentation of small lesions in breast cancer DCE-MRI images, making it difficult to achieve high-quality segmentation.

Method used

A multi-module fusion segmentation network - Harmonic Vision Mamba Framework (HV-Mamba), is proposed. By combining frequency domain feature extraction, spatial domain detail enhancement, attention mechanism and multi-scale convolutional unit, it integrates multi-dimensional Fourier transform, frequency spatial domain attention mechanism and adaptive selection convolution kernel to improve the accuracy and robustness of image segmentation.

Benefits of technology

It significantly improves the segmentation accuracy and stability of small lesions in breast cancer images, and is suitable for segmentation processing of complex backgrounds and subtle boundary features, providing efficient and accurate image segmentation support for early screening and clinical diagnosis of breast cancer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963545A_ABST
    Figure CN119963545A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-module fusion segmentation network implementation method for breast cancer DCE-MRI image segmentation. The method comprises the following steps: 1, collecting image data from a plurality of disclosed breast cancer DCE-MRI medical image data sets; 2, arranging and dividing the collected image data to obtain a training set and a test set; 3, preprocessing the image data in the step 2, and standardizing the data to ensure the consistency of different image devices or imaging conditions; 4, constructing a segmentation network HV-Mama combining a frequency domain and multi-scale attention, wherein the network comprises four core modules; 5, training the constructed segmentation network HV-Mama by using the data set preprocessed in the step 3; and 6, applying the weight obtained by training to a test set to evaluate the segmentation effect. Through independent processing and cooperative work of the modules, accurate segmentation of the breast tumor under the complex background is realized, and the wide application potential of the method in the field of medical image segmentation is shown.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image processing, and in particular to a method for implementing a multi-module fusion segmentation network for breast cancer DCE-MRI image segmentation. The method proposes an image segmentation network that combines frequency domain feature extraction, spatial domain detail enhancement, attention mechanism and multi-scale convolution unit, namely the Harmonic Vision Mamba Framework (HV-Mamba). HV-Mamba improves the segmentation accuracy and stability of small lesion areas in breast cancer images by integrating multidimensional Fourier transform (FFT), frequency-space domain attention mechanism, adaptive selection of convolution kernel and other technologies. The framework is suitable for the segmentation processing of complex background and subtle boundary features in breast cancer DCE-MRI, and provides efficient and accurate image segmentation support for early screening and clinical diagnosis of breast cancer. Background Art

[0002] In the field of medical imaging, accurate image segmentation plays a vital role in the diagnosis and treatment process. For example, in breast cancer detection, accurate lesion segmentation not only helps doctors to accurately evaluate the boundaries and volumes of the tumor area, but also supports the design of personalized treatment plans; in dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI), accurate distinction between tumor areas and normal tissues can help early diagnosis and significantly improve treatment outcomes. However, current image segmentation methods still lack accuracy and robustness when dealing with complex anatomical structures, artifacts, and noise interference, especially in the segmentation task of small lesion areas, where existing models find it difficult to achieve high-quality segmentation.

[0003] Currently, segmentation models based on encoder-decoder structures such as U-Net have been widely used in medical image processing. However, due to their limited receptive field, they are difficult to capture the long-range dependencies and multi-scale features of images, resulting in limited performance in high-noise and low-contrast breast images. At the same time, traditional frequency domain processing methods such as fast Fourier transform (FFT) can extract global frequency information of images and effectively deal with complex backgrounds, but they are limited in processing multi-scale and local edge details. In addition, although segmentation models based on graph convolutional networks can handle irregular geometric shapes, the stability of feature extraction and detail resolution are still not ideal in complex breast anatomical structures and artifact-rich images.

[0004] To address the above challenges, the present invention proposes a multi-module fusion segmentation network, HarmonicVision Mamba (HV-Mamba). HV-Mamba can capture key features in the frequency and spatial domains by introducing a frequency domain feature extraction module, a hybrid attention module that combines frequency and spatial domain attention, and an adaptive selection convolution unit and a spatial channel interaction module, significantly enhancing the ability to segment the global structure and subtle boundaries of breast cancer images. The modular design of HV-Mamba has shown significant improvements in dealing with complex background noise and edge detail processing, providing a high-precision and high-robustness solution for breast cancer detection and image segmentation. Summary of the invention

[0005] The existing segmentation methods have obvious limitations in dealing with complex noise, geometric complexity and edge feature ambiguity, especially in the segmentation of small lesion areas in breast cancer DCE-MRI images. The existing methods are difficult to meet the accuracy requirements. The present invention proposes a multi-module fusion segmentation network implementation method for breast cancer DCE-MRI image segmentation. The method proposes a multi-module fusion segmentation network - Harmonic Vision Mamba Framework (HV-Mamba) to solve the above problems in the prior art.

[0006] The technical solution included in the present invention to solve the technical problem includes the following steps: Step 1: Collect image data from multiple public breast cancer DCE-MRI medical imaging datasets; Step 2: Organize and divide the collected image data to obtain training sets and test sets; Step 3, preprocessing the image data in step 2, and standardizing the data to ensure consistency under different imaging devices or imaging conditions; Step 4: Construct a segmentation network HV-Mamba that combines frequency domain and multi-scale attention. The network includes four core modules: harmonic state space module HSS, hybrid coordinate frequency attention module HCF, self-selected convolution unit seSK, and spatial channel interaction module Spac; When the image data is input, the first branch will enter the harmonic state space module HSS, and the second branch will perform downsampling operations; the branch that enters the harmonic state space module HSS will output the obtained features through 4 residual operations, and the final output features and the results of the first three residuals will be input into a hybrid coordinate frequency attention module HCF respectively, and the features after the 4th residual operation will be input into the self-selective convolution unit seSK; the second branch will perform 4 downsampling operations, and the features obtained each time will be aligned with the features obtained after the residual operation of the first branch result. The features obtained by the two branches will be spliced ​​and input into the spatial channel interaction module Spac for channel shuffling to obtain the final result, and finally the target area will be generated through multiple convolution operations; Step 5: Use the preprocessed data set in step 3 to train the constructed segmentation network HV-Mamba; Step 6. Apply the trained weights to the test set to evaluate the segmentation effect. Tune the model through indicators such as Dice Similarity Coefficient (DSC), Mean Intersection over Union (mIoU), 95% Hausdorff Distance (HD95), Kappa coefficient (Kappa), and Matthew correlation coefficient (MCC) to further improve the segmentation accuracy and robustness.

[0007] Beneficial effects of the present invention: Based on the deep learning framework, the present invention takes into account the main problems currently faced by medical image processing, including noise interference, geometric complexity and semantic ambiguity, which lead to low segmentation accuracy and affect the diagnosis and treatment effects. The multi-aspect segmentation network proposed in the present invention significantly improves the accuracy and robustness of medical image segmentation by integrating multi-frequency feature modulation, deformation convolution and context analysis. Its innovative multi-dimensional feature fusion technology enhances the recognition ability of complex geometric structures and semantic information, and improves the generalization ability and applicability of the model through pre-training and refinement training. The application of this invention not only improves the accuracy of medical image segmentation, but also significantly improves the efficiency and effectiveness of clinical diagnosis and treatment, bringing important technological advances to the field of medical imaging, and ultimately improving patients' medical experience and treatment outcomes.

[0008] The network constructed by the present invention, in which the harmonic state space module HSS suppresses interference noise through multi-dimensional Fourier transform and global feature extraction; the hybrid coordinate frequency domain attention mechanism HCF uses spatial and frequency domain features to enhance the fine boundary expression; the self-selective convolution unit seSK selectively enhances local features through multi-scale convolution to improve the detection sensitivity of small lesions; the spatial channel fusion module SpaC integrates multi-domain features through the channel shuffling mechanism to balance the information flow. The present invention realizes the accurate segmentation of breast tumors under complex backgrounds through the independent processing and collaborative work of the above modules, demonstrating its wide application potential in the field of medical image segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 It is a schematic diagram of the overall network framework in an embodiment of the present invention.

[0010] Figure 2 Detailed structural diagram of Hybrid Coord-Freq Attention (HCF) in an embodiment of the present invention.

[0011] Figure 3 It is a detailed structural diagram of the self-selective convolution unit (SeSK) in an embodiment of the present invention.

[0012] Figure 4 It is a detailed structural diagram of the spatial channel interaction module (Spatio-Channel Confluent Module, SpaC) in an embodiment of the present invention. DETAILED DESCRIPTION

[0013] The present invention is further explained below with reference to the embodiments.

[0014] like Figure 1As shown, the present invention proposes a multi-module fusion segmentation network implementation method for breast cancer DCE-MRI image segmentation. HV-Mamba introduces four core modules to work together to improve the accuracy and robustness of breast cancer image segmentation. The network includes the following modules: Harmonic State Space Module (HSSM), which uses multidimensional Fourier transform and channel adjustment to capture long-range dependency information in breast cancer images and suppress noise; Hybrid Coord-Freq Attention Module (HCF), which combines spatial domain and frequency domain attention to accurately capture the detailed features of the image; Self-Selective Kernel Unit (seSK), which adaptively selects different convolution scales to effectively deal with the multi-scale lesion features of breast images; and Spatio-Channel Confluent Module (SpaC), which improves the segmentation performance of complex boundaries and detailed features in images through the interaction of channel and spatial features. The present invention shows significant performance advantages in breast cancer DCE-MRI segmentation, and has good scalability and practical application potential. This method designs a variety of modular components to extract global and local features of the image in the frequency domain and spatial domain, thereby improving the segmentation accuracy. Specifically, it includes the following steps: Step 1: Collect image data from multiple public breast cancer DCE-MRI medical imaging datasets, including BreastDM, I-SPY 1, and BCMedSet. Each dataset contains annotated tumor areas to ensure that the model has a comprehensive source of training and testing samples.

[0015] Step 2: Organize and divide the collected image data into training sets and test sets.

[0016] Step 3: Preprocess the image data in step 2. The data is standardized to ensure consistency across different imaging devices or imaging conditions, including uniform image size adjustment (e.g., adjusting the image size to ), and normalize the pixel values ​​to scope.

[0017] Step 4: In order to achieve high-precision medical image segmentation, a segmentation network called Harmonic Vision Mamba combining frequency domain and multi-scale attention is proposed. It includes four core modules: Harmonic State Space (HSS), Hybrid Coord-Freq (HCF), Self-Selective Kernel (seSK), and Spatio-Channel Confluent (Spac). Figure 1 In the framework shown, when the data image is input, one branch will enter the HSS module, and the other branch will perform downsampling operations. The branch that enters the HSS module for processing will output the obtained features through 4 residual operations. At the same time, the final output features and the results of the first three residuals will be input into an HCF module respectively, and the features of the 4th residual operation will be input into the seKS unit. The other branch will perform 4 downsampling operations, and the features obtained each time will be aligned with the features obtained after the residual operation of the other branch result. The features obtained from the two branches will be spliced ​​and input into the Spac module for channel shuffling to obtain the final result, and finally the target area will be generated through multiple convolution operations.

[0018] The following is a detailed description of the structure and function of each encoder, and the relevant mathematical formulas are given.

[0019] In the harmonic state space module HSS module, the input image is first subjected to a two-dimensional Fourier transform: (1) in, represents the input data in the spatial domain, is the output after conversion to the frequency domain, M and N are the width and height of the input image respectively, and are the frequency coordinates in the horizontal and vertical directions; the real and imaginary parts of the frequency domain output are separated and connected in the channel dimension to form a frequency domain feature representation containing amplitude and phase information: (2) in, and denote the real and imaginary parts of the frequency domain output, respectively, and are used to capture the amplitude and phase characteristics.

[0020] Next, the harmonic state space module HSS module performs multi-directional scanning through its own vision self-similarity block (Vision Self-Similarity Block, VSS block) to divide the frequency domain features into small blocks in order to retain local frequency information and capture global dependencies; for the size of The input image is segmented to produce sub-regions, thereby achieving multi-directional feature extraction: (3) in represents the block operator, Represents a single block.

[0021] Then, the dynamic time constant Adjust the processing scale to make the model flexible to different feature scales: (4) in, is a random factor, and is the maximum and minimum range of the dynamic time constant. The features are passed through the projection matrix Perform a direction-specific linear transformation: (5) in, Indicates direction The scan feature output on is the corresponding projection matrix. The final output feature is obtained by weighted aggregation of the scanning features in each direction : (6) Through multi-directional processing, dynamic time constant adjustment and projection integration, the characteristics It achieves the capture of global complex structures and adaptation to local details.

[0022] Finally, the frequency domain representation is restored to the spatial domain through inverse Fourier transform, so as to integrate the global and local frequency domain features in the spatial domain: (7) in, In order to convert back to the feature representation in the spatial domain, the fusion of frequency domain and spatial domain features is completed, thereby effectively improving the global and detail accuracy of image segmentation.

[0023] like Figure 1 and Figure 3As shown in the figure, the hybrid coordinate frequency attention module HCF performs fast Fourier transform on the input data, and enhances the expression of details in the frequency domain and spatial domain by applying the coordinate attention mechanism to the real and imaginary parts. Then, the composite features are restored to the spatial domain through convolution and inverse Fourier transform to improve the clarity of edge and texture features. The specific implementation is as follows: First, the spatial domain is represented As input features and converted to the frequency domain, the lesion details are extracted through predefined frequency bands while attenuating noise interference. Its mathematical expression is: (8) in, represents the time domain input signal sequence, is the frequency domain representation, indicating the frequency index The amplitude of the signal component at is the extracted frequency component.

[0024] Then, the coordinate attention is applied to the real and imaginary parts of one dimension respectively, denoted as and Among them, the real amplitude captures the macroscopic structural changes, the coordinate attention enhances the spatially correlated lesion signals and suppresses the background noise; and the phase information of the imaginary part helps to distinguish spatially similar but structurally different regions. The attention mechanism of each component is defined as follows: (9) (10) in, represents the convolution weight in the coordinate attention mechanism, is the Sigmoid activation function; , denote the real and imaginary parts of the initial coordinates after attention is applied, respectively.

[0025] After the initial coordinate attention is applied, the real and imaginary parts are iteratively refined through a feedback loop, with convolutional layers and sigmoid activations to enhance sensitivity to small lesions and suppress background noise. The iterative process is: (11) (12) in, is the attention map, Conv represents the convolution operation. After multiple iterations, the real and imaginary parts are reconstructed into complex form to obtain complex features : (13) The plural feature Then iFFT is converted back to the spatial domain to realize the reintegration of frequency domain features with spatial domain. The expression is: (14) Among them, through this attention mechanism, the model realizes frequency domain optimization in the spatial domain, ensuring high accuracy of lesion localization and reducing noise interference.

[0026] like Figure 4 As shown in the figure, in the self-selective convolution unit seSK, the global context of the input features obtained after the residual operation is first encoded through the self-attention mechanism to establish associations between pixels. The self-attention mechanism contains three key components: query (Q), key (K), and value (V). Among them, Q Represents pixel or area features and is used to identify relevant features in an image; K Indicates the importance of each pixel or area, helping with alignment Q ; V contains the actual content and generates the output through weighting and aggregation. Input features First pass The volume integral is decomposed into Q , K , V Tensor: (15) Tensor through Scaling to stabilize values, computing global dependencies: (16) After attention weighting, the features are obtained Then, multi-scale convolution kernels are applied to the weighted features, and each kernel size generates a unique scale representation. Represents multi-scale output: (17) in represents the convolution kernel, Represents the attention-weighted features. Large convolution kernels capture broad structural dependencies, while small convolution kernels focus on local changes. All multi-scale outputs are summed element-wise to obtain a unified feature tensor : (18) after, Perform global pooling in the spatial dimension to obtain the channel statistics vector : (19) The fully connected layer will count the vector Mapping to scale-specific weights : (20) Among them, the weight After Softmax regularization, the scale modulation parameters are generated , get the adjusted multi-scale fusion features : (twenty one) Therefore, the self-selective convolutional unit seSK dynamically prioritizes scale-related features, facilitating accurate lesion segmentation in both global and local contexts.

[0027] Figure 4 The structure diagram of the spatial channel fusion module SpaC is shown, which processes the features of the harmonic state space module HSS and the spatial encoder path through a channel shuffle operation. This operation breaks the isolated grouping of channels, thereby achieving uniform propagation of cross-domain features: (twenty two) Among them, among them, and denote the features from the harmonic state space module HSS and the spatial encoder path, respectively, represents the channel shuffle operation, Indicates channel splicing, is the number of groups, is the output after shuffling.

[0028] Next, the spatial dimension is compressed through maximum pooling and average pooling, and the maximum and mean values ​​of the channels are extracted to generate a dynamic channel attention map: (twenty three) in, is the Sigmoid activation function, and are weights and biases respectively, and It is the maximum pooling and average pooling operation.

[0029] To further improve the spatial feature expression, the spatial channel fusion module SpaC introduces spatial attention. Convolution (Conv1) on input features Dimensionality reduction , after activation , and then convolution (Conv2) generates a spatial attention map. The final generated spatial attention map is Multiply element by element to get , to highlight key spatial areas: (twenty four) Therefore, the SpaC module combines channel and spatial attention to refine boundary expression and achieve a balance between frequency and spatial domain features, thereby supporting the fine characterization of tumor boundaries in complex regions.

[0030] Step 5: Train HV-Mamba using the preprocessed and data augmented dataset from step 3. The framework is implemented using PyTorch and trained on dual NVIDIA 3080 GPUs using specific hyperparameters and data augmentation techniques.

[0031] Multiple loss functions are used in the training process, such as cross entropy loss and Dice similarity coefficient loss, to ensure the high accuracy and robustness of the model in lesion segmentation. The performance of the model is evaluated on three benchmark datasets: BreastDM, I-SPY 1, and BCMedSet.

[0032] Step 6: Apply the trained weights to the test set to evaluate the segmentation effect. We use the main and key evaluation indicators to evaluate the model, including 95% Hausdorff distance (HD95), Dice Similarity Coefficient (DSC), geometric mean (G-mean), Kappa coefficient (Kappa), and Matthew correlation coefficient (MCC). HD95 is used to measure the maximum and minimum distance between two point sets. For a given two point sets, and , Hausdorff distance is defined as: (25) in, Indicate point and Point In order to reduce the influence of noise and outliers, the Euclidean distance between HD95 is defined as: (26) in (27) DSC is an indicator to measure the similarity between two sample sets. The calculation formula is: (28) in and Represent the sizes of the two sample sets, Represents their intersection. The value range of DSC is from 0 to 1, and the larger the value, the higher the similarity.

[0033] G-Mean is a comprehensive indicator of classifier performance, especially suitable for dealing with class imbalance problems. Its calculation formula is: (29) Among them, Sensitivity indicates the proportion of correctly identified positive examples, and Specificity indicates the proportion of correctly identified negative examples. The higher the G-Mean value, the better the overall performance of the classifier when processing unbalanced data sets.

[0034] The Kappa coefficient is used to measure the consistency between the classification results of the classifier and the random classification results. The calculation formula is: (30) in, represents the consistency ratio of observations, Indicates the random consistency ratio. The Kappa coefficient ranges from -1 to 1, with higher values ​​indicating better consistency and 0 indicating consistency with random classification results.

[0035] MCC is a balanced classification performance metric that takes into account true positives, false positives, true negatives, and false negatives. Its calculation formula is: (31) Among them, TP, TN, FP and FN represent true positive, true negative, false positive and false negative respectively. The value range of MCC is The higher the value, the better the classification performance. It is a more comprehensive indicator than accuracy.

[0036] According to the experimental results, the multi-module fusion segmentation network framework of the present invention shows the best performance in all indicators of almost all categories, indicating that it has significant advantages and strong competitiveness in the breast cancer DCE-MRI image segmentation task. This result further verifies the effectiveness and robustness of the multi-module fusion network in processing complex medical image segmentation tasks.

[0037] Table 1 Comparison of experimental results on DCE-MRI dataset

[0038] The quantitative results of the DCE-MRI dataset are shown in Table 1. The results show that the HV-Mamba (ours) proposed in this paper performs well in multiple indicators. The quantitative results of the I-SPY 1 dataset are shown in Table 2.

[0039] Table 2 Comparison of experimental results on the I-SPY 1 dataset

[0040] Table 3 Comparison of experimental results of BCMedSet dataset

[0041] In order to verify the impact of each module on the network, we evaluated the impact of each key module in HV-Mamba on the overall network performance, and the experimental results are shown in Table 4. By ablating HSS Module, HCF Attention, seSK Unit and SpaCModule one by one, we can quantify the contribution of each module to the model.

[0042] Table 4 Ablation experiments on key modules of Harmonic Vision Mamba on the BreastDM dataset. √ means retaining the module, and the unchecked position means deleting the module from the network.

[0043]

[0044] Experiments show that the model performs best when all modules work together. In the ablation experiments of each module, the removal of HSSModule has the greatest impact on the performance of the model, which shows that the dual-channel processing method of HSS Module plays a key role in capturing and coordinating global contextual information. The absence of HCF Attention caused a large decrease in G-mean and an increase in HD95 from 1.76 to 1.8, which means that the multi-scale feature fusion and the interaction between feature channels in HCF Attention play an important role in detecting lesions and boundary accuracy. The removal of the SpaC module reduced Dice to 74.28% and increased HD95 to 1.81.

[0045] In summary, we proposed a framework called Harmonic Vision Mamba (HV-Mamba), which combines multi-dimensional FFT and the attention mechanism in the frequency and spatial domains to introduce multi-scale convolution kernels, coordinate the balanced flow of features in the frequency and spatial domains, generate information-rich feature representations, and improve the segmentation accuracy of tumors in breast cancer DCE-MRI. Among them, HSSModule integrates frequency and spatial domain information to obtain global structural information. HCF Attention uses the joint attention mechanism in the coordinate and frequency domains to capture the local heterogeneity and complex morphological characteristics of breast cancer lesions. seSK Unit adaptively selects the most suitable convolution kernel to process tumor structures of different scales and enhances the model's adaptability to multi-scale features. SpaC Module uses the fusion processing of spatial and channel information to improve the diversity and richness of feature expression, providing more powerful support for the accurate segmentation of small tumor areas. Experiments show that the HV-Mamba framework shows the current best performance in the segmentation of small tumors in breast cancer, and together with reconstruction, it provides important support for the early detection and diagnosis of breast cancer in clinical practice.

[0046] The above contents are only preferred implementation modes of the present invention and do not limit the protection scope of the present invention. It should be emphasized that any minor adjustments and improvements made by technicians in this field without departing from the basic principles of the present invention should be regarded as the protection scope of the present invention.

Claims

1. A multi-module fusion segmentation network implementation method for breast cancer DCE-MRI image segmentation, characterized in that: The steps include: Step 1: Collect image data from multiple public breast cancer DCE-MRI medical imaging datasets; Step 2: Organize and divide the collected image data to obtain training sets and test sets; Step 3, preprocessing the image data in step 2, and standardizing the data to ensure consistency under different imaging devices or imaging conditions; Step 4: Construct a segmentation network HV-Mamba that combines frequency domain and multi-scale attention. The network includes four core modules: harmonic state space module HSS, hybrid coordinate frequency attention module HCF, self-selected convolution unit seSK, and spatial channel interaction module Spac; When the image data is input, the first branch will enter the harmonic state space module HSS, and the second branch will perform downsampling operations; the branch that enters the harmonic state space module HSS will output the obtained features through 4 residual operations, and the final output features and the results of the first three residuals will be input into a hybrid coordinate frequency attention module HCF respectively, and the features after the 4th residual operation will be input into the self-selective convolution unit seSK; the second branch will perform 4 downsampling operations, and the features obtained each time will be aligned with the features obtained after the residual operation of the first branch result. The features obtained by the two branches will be spliced ​​and input into the spatial channel interaction module Spac for channel shuffling to obtain the final result, and finally the target area will be generated through multiple convolution operations; Step 5: Use the preprocessed data set in step 3 to train the constructed segmentation network HV-Mamba; Step 6: Apply the trained weights to the test set to evaluate the segmentation effect.

2. The method for implementing a multi-module fusion segmentation network for breast cancer DCE-MRI image segmentation according to claim 1, characterized in that: The specific steps of the harmonic state space module HSS are as follows: First, the harmonic state space module HSS converts the input spatial domain image data into a frequency domain representation through a two-dimensional Fourier transform to separate global features and local noise components: (1) in, represents the input data in the spatial domain, is the output after conversion to the frequency domain, M and N are the width and height of the input image respectively, and are the frequency coordinates in the horizontal and vertical directions; the real and imaginary parts of the frequency domain output are separated and connected in the channel dimension to form a frequency domain feature representation containing amplitude and phase information: (2) in, and denote the real and imaginary parts, respectively, and are used to capture the amplitude and phase characteristics; Subsequently, the harmonic state space module HSS performs multi-directional scanning through the visual state space VSS block to split the frequency domain features into small blocks in order to preserve local frequency information and capture global dependencies; for The input image is segmented to produce sub-regions, thereby achieving multi-directional feature extraction: (3) in represents the block operator, Represents a single block; Next, the harmonic state space module HSS introduces a dynamic time constant Adjust the processing scale so that the model can adapt to multiple feature scales: (4) in, is a random factor, and is the maximum and minimum range of the dynamic time constant; each block The features are passed through the projection matrix Perform a direction-specific linear transformation: (5) in, Indicates direction The scan feature output on is the corresponding projection matrix; the final output feature is obtained by weighted aggregation of the scanning features in each direction : (6) Through multi-directional processing, dynamic time constant adjustment and projection integration, the characteristics It realizes the capture of global complex structures and the adaptation of local details; Finally, the frequency domain features are restored to the spatial domain through inverse Fourier transform to achieve the integration of global and local features: (7) in, In order to convert back to the feature representation in the spatial domain, the fusion of frequency domain and spatial domain features is completed.

3. The method for implementing a multi-module fusion segmentation network for breast cancer DCE-MRI image segmentation according to claim 2, characterized in that: The hybrid coordinate frequency attention module HCF performs fast Fourier transform on the input data, and enhances the detail expression in the frequency domain and spatial domain by applying the coordinate attention mechanism to the real and imaginary parts, and then restores the composite features to the spatial domain through convolution and inverse Fourier transform to improve the clarity of edge and texture features. The specific implementation is as follows: First, the spatial domain is represented As input features and converted to the frequency domain, the lesion details are extracted through predefined frequency bands while attenuating noise interference. Its mathematical expression is: (8) in, represents the time domain input signal sequence, is the frequency domain representation, indicating the frequency index The amplitude of the signal component at is the extracted frequency component; Then, coordinate attention is applied to the real part of the one-dimensional and the imaginary part , the attention mechanism of each component is defined as follows: (9) (10) in, represents the convolution weight in the coordinate attention mechanism, is the Sigmoid activation function, , denote the real and imaginary parts of the initial coordinates after attention is applied, respectively; After the initial coordinate attention is applied, the real and imaginary parts are iteratively refined through a feedback loop, with convolutional layers and sigmoid activations to enhance sensitivity to small lesions and suppress background noise. The iterative process is: (11) (12) in, is the attention map, Conv represents the convolution operation; after multiple iterations, the real and imaginary parts are reconstructed into complex form to obtain complex features : (13) The plural feature Then iFFT is converted back to the spatial domain to realize the reintegration of frequency domain features with spatial domain. The expression is: (14)。 4. The method for implementing a multi-module fusion segmentation network for breast cancer DCE-MRI image segmentation according to claim 3, characterized in that: The self-selected convolution unit seSK is specifically implemented as follows: First, the global context of the input feature obtained after the residual operation is encoded through the self-attention mechanism; the input feature First pass The volume integral is decomposed into Q , K , V Tensor: (15) Among them, the tensor through Scaling to stabilize values, computing global dependencies: (16) Next, the multi-scale convolution kernel is applied to the weighted features , each convolution kernel size generates a unique scale representation, using Represents multi-scale output: (17) in, Represents convolution kernels of different scales, Represents the features obtained after attention weighting; all multi-scale outputs are summed element by element to obtain a unified feature tensor : (18) Then, the unified feature tensor Perform global pooling in the spatial dimension to obtain the channel statistics vector : (19) The fully connected layer FC transforms the channel statistics vector Mapping to scale-specific weights : (20) Finally, the weight After Softmax regularization, the scale modulation parameters are generated , get the adjusted multi-scale fusion features : (21)。 5. The method for implementing a multi-module fusion segmentation network for breast cancer DCE-MRI image segmentation according to claim 4, characterized in that: The spatial channel fusion module SpaC is specifically implemented as follows: The spatial channel fusion module SpaC processes the features of the harmonic state space module HSS and the spatial encoder path through a channel shuffle operation, breaking the isolated grouping of channels and thus achieving uniform propagation of cross-domain features: (22) in, and denote the features from the harmonic state space module HSS and the spatial encoder path, respectively, represents the channel shuffle operation, Indicates channel splicing, is the number of groups, is the output after shuffling; Next, the spatial dimension is compressed through maximum pooling and average pooling, and the maximum and mean values ​​of the channels are extracted to generate a dynamic channel attention map: (23) in, is the Sigmoid activation function, and are weights and biases respectively, and It is the maximum pooling and average pooling operation; Finally, spatial attention is introduced into the spatial channel fusion module SpaC. Convolutional channel attention map of the input Dimensionality reduction , after activation , and then convolution is performed to generate the spatial attention map. Finally, the spatial attention map and the channel attention map Multiply element by element to get , to highlight key spatial areas: (24)。

Citation Information

Patent Citations

  • Underwater image enhancement method of Mama hybrid architecture based on space-frequency fusion

    CN118710507A

  • Mammary gland ultrasound image segmentation network based on diffusion probability model

    CN118864493A

  • Multi-aspect segmentation network implementation method for visible light medical image segmentation

    CN119180963A

  • System and method for adverse event detection or severity estimation from surgical data

    US20200265273A1

  • Scattering vision transformer

    US20240386238A1

Cited By

  • Remote sensing image building extraction method and system based on visual Mama model

    CN120599504A

  • Cross-modal breast image segmentation and classification method based on X-ray and clinical medicine text

    CN121259327A

  • Multi-region fusion classification method and system based on breast magnetic resonance image

    CN121415163A

  • A multi-region fusion classification method and system based on breast magnetic resonance images

    CN121415163B