Breast tumor segmentation method of multi-modal multi-stage feature fusion learning network
Through the multimodal multi-level feature fusion learning network, the multi-scale features of MRI images are deeply integrated, which solves the problem of poor segmentation accuracy in breast tumor segmentation, and achieves higher accuracy and robustness.
Patent Information
- Application Number
- CN202411909903.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art has poor segmentation accuracy in breast tumor segmentation, mainly due to the blurred tumor boundaries and complex morphology in MRI images, the traditional convolutional neural network model is difficult to accurately capture, and it is insufficiently sensitive to morphological characteristics, so it is easy to segment incorrectly.
A multimodal multi-level feature fusion learning network is adopted, and pre-processed through four single-modal MRI images (DCE, DWI, T1 weighted, and T2 weighted images) is used to enhance morphological and edge features, and combined with the Transformer encoder module and feature fusion subnet, multi-scale features are deeply fused to improve the neural network's ability to mine and associate modal features.
It significantly improves the understanding and detection ability of breast tumors by neural networks, effectively resists the interference of vague and irregular boundaries of tumor lesion areas, improves the accuracy and robustness of breast tumor segmentation, and enhances generalization ability.
Smart Images

Figure CN119992079A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical image processing, and in particular to a breast tumor segmentation method of a multi-modal multi-level feature fusion learning network. Background Art
[0002] In recent years, the incidence and mortality of breast cancer have shown an increasing trend year by year. Since Magnetic Resonance Imaging (MRI) can provide doctors with three-dimensional visualization data of human tissues and organs and has a high soft tissue contrast, MRI has become a common auxiliary tool for clinical diagnosis and analysis of tumors. Doctors can accurately determine the location, size and morphological characteristics of tumors through MRI, and can evaluate the malignancy of tumors and predict treatment response. Therefore, MRI plays an important role in the screening, diagnosis, staging and follow-up of breast tumors. Therefore, computer-assisted segmentation of breast tissue based on MRI has important clinical value for improving the diagnosis and treatment efficiency of breast tumors.
[0003] With the widespread application of machine learning and deep learning technologies in the field of medical image processing, a variety of breast tumor segmentation methods based on deep learning have emerged. Wang et al. used extreme learning machines to extract and classify breast MRI images, and then detected the tumor area in the image. However, this method has obvious limitations in extracting tumor features with blurred edges or complex morphology, and the accuracy of the segmentation results is insufficient; Sun et al. proposed a medical image segmentation method based on unsupervised domain adaptation using orthogonal decomposition technology, decomposing image features of different imaging modes to improve segmentation accuracy. However, the model has insufficient feature association ability for multimodal images and cannot fully utilize multimodal information in breast tumor segmentation tasks, which affects its segmentation performance; Zunair et al. proposed a breast MRI image segmentation method based on Sharp U-Net, which reduces the number of parameters through deep separable convolution, and improves the segmentation accuracy and computational efficiency of medical images to a certain extent; Alanarasu et al. combined multilayer perceptrons with MRI medical images and proposed a UNeXt neural network based on multilayer perceptrons, which can achieve fast and efficient breast tumor image segmentation. Although Sharp U-Net and UNeXt can generate multi-scale feature maps for MRI images of breast tumors, their feature sampling process and fusion strategy are simply designed and do not perform well when dealing with tumor segmentation tasks of different shapes and sizes.
[0004] In general, the main reason for the poor segmentation accuracy of existing technologies is that, on the one hand, when performing breast tumor segmentation, the boundaries of breast tumor lesion areas in MRI images are usually fuzzy, irregular and have certain changes, which may cause the traditional convolutional neural network model to be unable to accurately capture and segment these complex boundaries, and due to the limitation of the local receptive field, the segmentation results are not fine enough. On the other hand, although the MRI images of breast tissue contain rich detailed information, most of the existing technologies ignore the enhancement of the unique characteristics of breast tumors, which leads to the segmentation model of tumor MRI images often lacking sensitivity to morphological features, and easily mis-segmenting tumor tissue with fuzzy boundaries as normal tissue. In addition, since medical image datasets are usually small and the annotation cost is high, when the training data is not enough or not diverse enough, the traditional convolutional neural model is prone to overfitting, so that the existing methods perform well on the training set, but have poor generalization ability on the test set. Summary of the invention
[0005] The present invention aims to solve the above-mentioned technical problems existing in the prior art and provides a breast tumor segmentation method based on a multi-modal multi-level feature fusion learning network.
[0006] The technical solution of the present invention is: a breast tumor segmentation method of a multi-modal multi-level feature fusion learning network, which is carried out according to the following steps:
[0007] Step 1. Input any number of four single-modality breast tumor MRI images to form a training set T, wherein the four single-modality MRI images include DCE images, DWI images, T1-weighted images, and T2-weighted images;
[0008] Step 2: Preprocess each breast tumor MRI image I in the training set to obtain an image training set T with enhanced morphological features. morph And edge feature enhanced image training set T edge , let the height of image I be H and the width be W;
[0009] Step 2.1: Use the convolution kernel H shown in formula (1) to perform convolution operation on image I to enhance the contrast of the edge of the breast tumor lesion in image I, thereby obtaining image I with enhanced morphological features. 1 , forming the image training set T with enhanced morphological features morph ;
[0010]
[0011] Step 2.2 Let the value of the pixel at coordinate (x, y) in image I be f(x, y), and use the gradient difference method to calculate the horizontal gradient value Grad of f(x, y) respectively. x(x, y), vertical gradient value Grad y (x, y), main diagonal gradient value Grad d1 (x, y), sub-diagonal gradient value Grad d2 (x, y), and then calculate the weighted gradient value Grad(x, y) of f(x, y) according to formula (2) to obtain the image I after edge feature enhancement 2 , forming the edge feature enhanced image training set T edge ;
[0012] Grad(x,y)=ω 1 ×Grad x (x,y)+ω 2 ×Grad y (x,y)+ω 3 ×Grad d1 (x,y)+ω 4 ×Grad d2 (x,y) (2)
[0013] 1≤x≤W, 1≤y≤H, ω 1 ,ω 2 ,ω 3 and ω 4 is the preset weight constant;
[0014] Step 3. Establish and initialize the multimodal multi-level feature fusion learning network N seg , including 1 morphological feature extraction sub-network N morph , 1 edge feature extraction sub-network N edge , 1 feature fusion sub-network N fusion , 1 segmentation sub-network N segHead ;
[0015] Step 3.1 Establish and initialize the morphological feature extraction subnetwork N morph , including 3 Transformer encoder modules, namely Said The input data dimension is W×H, and the output data dimension is W / 4×H / 4×C 1 , The input data dimension is W / 4×H / 4×C 1 , the output data dimension is W / 8×H / 8×C 2 , The input data dimension is W / 8×H / 8×C 2 , the output data dimension is W / 16×H / 16×C 3 , C 1 , C 2 and C 3 is a preset constant;
[0016] Step 3.2 Establish and initialize the edge feature extraction subnetwork N edge , including 3 Transformer encoder modules, namely Said The input data dimension is W×H, and the output data dimension is W / 4×H / 4×C 1 , The input data dimension is W / 4×H / 4×C 1 , the output data dimension is W / 8×H / 8×C 2 , The input data dimension is W / 8×H / 8×C 2 , the output data dimension is W / 16×H / 16×C 3 ;
[0017] Step 3.3 Establish and initialize the feature fusion subnetwork N fusion , including a cross-modal feature enhancement module M cfe , 1 Transformer encoder module 1 cross-modal feature enhancement fusion module M cffu , the cross-modal feature enhancement module M cfe Contains 1 Sigmoid activation function layer, 1 convolution layer with a convolution kernel size of 3×3, 1 batch normalization layer, and 1 ReLU activation function layer;
[0018] Step 3.4 Establish and initialize the segmentation subnetwork N segHead , including a sequence decoder module M Decoder , the M Decoder Contains 1 upsampling layer, 1 transposed convolution layer, 1 ReLU activation function layer, 1 batch normalization layer, and 1 Softmax activation function layer;
[0019] Step 4. Input the image training set T with enhanced morphological features morph And edge feature enhanced image training set T edge , for the multi-modal multi-level feature fusion learning network N seg Conduct training;
[0020] Step 4.1 T morph Each image I 1 Input morphological feature extraction subnetwork N morph to process;
[0021] Step 4.1.1 Image I 1 Divide into (W×H) / (4×4) image blocks, each image block is 4×4 in size, and get an image block sequence
[0022] Step 4.1.2 Using the Transformer encoder module right Processing is performed to obtain the morphological feature map at the first scale
[0023] Step 4.1.3 Using the Transformer encoder module right Processing is performed to obtain the morphological feature map at the second scale
[0024] Step 4.1.4 Using the Transformer encoder module right Processing is performed to obtain the morphological feature map at the third scale
[0025] Step 4.2 T edge Each image I 2 Input edge feature extraction subnetwork N edge to process;
[0026] Step 4.2.1 Image I 2 Divide into (W×H) / (4×4) image blocks, each image block is 4×4 in size, and get an image block sequence
[0027] Step 4.2.2 Using the Transformer encoder module right Processing is performed to obtain the edge feature map at the first scale
[0028] Step 4.2.3 Using the Transformer encoder module right Processing is performed to obtain the edge feature map at the second scale
[0029] Step 4.2.4 Using the Transformer encoder module right Processing is performed to obtain the edge feature map at the third scale
[0030] Step 4.3: Feature maps at different scales and Input feature fusion subnetwork N fusion to process;
[0031] Step 4.3.1 Utilize cross-modal feature enhancement module Mcfe , according to formula (3) to formula (5), calculate the feature weight λ at different scales 1 , 2 and λ 3 ;
[0032]
[0033] The lambda 1 represents the feature weight at the first scale, λ 2 represents the feature weight at the second scale, λ 3 represents the feature weight at the third scale, σ(·) represents the Sigmoid activation function, Conv T (·) represents a series of operations consisting of a convolution layer with a convolution kernel size of 3×3, a batch normalization layer, and a ReLU activation function layer. Concat(·) represents a concatenation operation.
[0034] Step 4.3.2 According to formula (6) and formula (7), the feature map and Perform weighted operation to obtain the weighted feature map and
[0035]
[0036] i∈{1,2,3},λ i represents the feature weight at the i-th scale, represents the morphological feature map at the i-th scale, represents the edge feature map at the i-th scale, represents the weighted morphological feature map at the i-th scale, Represents the weighted edge feature map at the i-th scale;
[0037] Step 4.3.3 Using the Transformer encoder module According to formula (8), the morphological feature maps and edge feature maps at different scales are processed to obtain the fused feature map and
[0038]
[0039] Said represents the fusion feature map at the i-th scale, They represent the fused feature map at the first scale, the fused feature map at the second scale, and the fused feature map at the third scale respectively;
[0040] Step 4.3.4 Enhance the fusion module M using cross-modal features cffu According to formula (9), the unweighted edge feature map, morphological feature map and fusion feature map are connected to obtain the feature map Then according to formula (10), we get the fusion feature map
[0041]
[0042] Said represents the fused connection feature map at the i-th scale, represents the fusion feature map at the i-th scale, represents the matrix Hadamard product operation;
[0043] Step 4.4 Use the segmentation subnetwork N segHead Fusion feature map Processing is performed to obtain the final segmentation result;
[0044] Step 4.4.1: Calculate the upsampled mapping feature map according to formula (11) and formula (12);
[0045]
[0046] And order Said represents the upsampled mapping feature map at the first scale, represents the upsampled mapping feature map at the second scale, Represents the upsampled mapping feature map at the third scale;
[0047] Step 4.4.2: According to formula (13), the upsampled mapping feature map is connected with the final fusion feature map;
[0048]
[0049] Said represents the composite connection feature map at the i-th scale, Represents the upsampled mapping feature map at the i-th scale;
[0050] Step 4.4.3 Composite connection feature graph The convolution operation with a kernel size of 3×3, the ReLU activation operation, and the batch normalization operation are performed in sequence to obtain the output feature map O i , the O i Represents the output feature map at the i-th scale;
[0051] Step 4.4.4: transform the feature map O iInput the Softmax activation function to obtain the category probability of each pixel and the lesion segmentation result Figure I segment And the trained multi-modal multi-level feature fusion learning network N seg ;
[0052] Step 5. Input the breast tumor MRI image J to be processed and use the trained deep convolutional neural network N seg Process J and output the lesion segmentation result image J output .
[0053] Compared with the prior art, the present invention has the following advantages: First, by deeply fusing the features of MRI images of four modalities at three different scales, the present invention can significantly improve the mining and association ability of the neural network for modal features. When processing complex tumor lesion areas, this multi-modal, multi-level feature fusion mechanism helps to improve the neural network's understanding and detection ability of breast tumors, thereby effectively resisting the interference of fuzzy and irregular boundaries in tumor lesion areas. Second, a specific independent module for feature enhancement and fusion is designed to fully capture the subtle structure and complex shape of breast tumors, reduce the loss of spatial information, and thus enhance the neural network's perception of tumors of different sizes and shapes, and improve the accuracy and robustness of breast tumor segmentation. Third, in the learning process, the mechanism of extracting features layer by layer can generate multi-scale features of tumor lesion areas. The resolution of the feature map gradually decreases as the network deepens, while the number of channels gradually increases as the network deepens, which is conducive to taking into account the extraction of local and global features of tumor lesion areas and ensuring the generalization ability of the neural network. In summary, the invention has the characteristics of strong morphological feature perception ability, excellent modal feature mining association performance, high accuracy, good robustness, and strong generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 This is a diagram showing the enhanced results of the morphological features of a breast tumor lesion according to an embodiment of the present invention.
[0055] Figure 2 This is a diagram showing the edge feature enhancement result of a breast tumor lesion according to an embodiment of the present invention.
[0056] Figure 3 4 is a diagram of a neural network structure according to an embodiment of the present invention.
[0057] Figure 4 This is a comparison chart of breast tumor segmentation results between the method of the present invention and a typical method. DETAILED DESCRIPTION
[0058] The present invention provides a breast tumor segmentation method using a multi-modal multi-level feature fusion learning network, which is characterized by being performed in the following steps:
[0059] Step 1. Input any number of four single-modality breast tumor MRI images to form a training set T, wherein the four single-modality MRI images include FLAIR images, T1-weighted images, T2-weighted images, and T1-Contrast images;
[0060] Step 2: Preprocess each breast tumor MRI image I in the training set to obtain an image training set T with enhanced morphological features. morph And edge feature enhanced image training set T edge , let the height of image I be H and the width be W;
[0061] Step 2.1 Use the convolution kernel H shown in formula (1) to perform convolution operation on image I to enhance the contrast of the edge of the breast tumor lesion in image I, so as to obtain Figure 1 The image I shown is after the morphological features are enhanced 1 ;
[0062]
[0063] Step 2.2 Let the value of the pixel at coordinate (x, y) in image I be f(x, y), and use the gradient difference method to calculate the horizontal gradient value Grad of f(x, y) respectively. x (x, y), vertical gradient value Grad y (x, y), main diagonal gradient value Grad d1 (x, y), sub-diagonal gradient value Grad d2 (x, y), and then calculate the weighted gradient value Grad(x, y) of f(x, y) according to formula (2), and get Figure 2 The image I shown is after edge feature enhancement 2 ;
[0064] Grad(x,y)=ω 1 ×Grad x (x,y)+ω 2 ×Grad y (x,y)+ω 3 ×Grad d1 (x,y)+ω 4 ×Grad d2 (x,y) (2)
[0065] 1≤x≤W, 1≤y≤H, ω 1 ,ω 2 ,ω 3 and ω 4 is a preset weight constant. In the embodiment of the present invention, let ω 1 =0.4,ω 2 =0.4,ω 3=0.3,ω 4 =0.2;
[0066] Step 3. Establish and initialize the multimodal multi-level feature fusion learning network N seg , including 1 morphological feature extraction sub-network N morph , 1 edge feature extraction sub-network N edge , 1 feature fusion sub-network N fusion , 1 segmentation sub-network N segHead ;
[0067] Step 3.1 Establish and initialize the morphological feature extraction subnetwork N morph , including 3 Transformer encoder modules, namely Said The input data dimension is W×H, and the output data dimension is W / 4×H / 4×C 1 , The input data dimension is W / 4×H / 4×C 1 , the output data dimension is W / 8×H / 8×C 2 , The input data dimension is W / 8×H / 8×C 2 , the output data dimension is W / 16×H / 16×C 3 , C 1 , C 2 and C 3 is a preset constant. In the embodiment of the present invention, let C 1 =256, C 2 =512, C 3 =1024;
[0068] Step 3.2 Create and initialize Figure 3 The edge feature extraction subnetwork N shown edge , including 3 Transformer encoder modules, namely Said The input data dimension is W×H, and the output data dimension is W / 4×H / 4×C 1 , The input data dimension is W / 4×H / 4×C 1 , the output data dimension is W / 8×H / 8×C 2 , The input data dimension is W / 8×H / 8×C 2 , the output data dimension is W / 16×H / 16×C 3 ;
[0069] Step 3.3 Establish and initialize the feature fusion subnetwork N fusion, including a cross-modal feature enhancement module M cfe , 1 Transformer encoder module 1 cross-modal feature enhancement fusion module M cffu , the cross-modal feature enhancement module M cfe Contains 1 Sigmoid activation function layer, 1 convolution layer with a convolution kernel size of 3×3, 1 batch normalization layer, and 1 ReLU activation function layer;
[0070] Step 3.4 Establish and initialize the segmentation subnetwork N segHead , including a sequence decoder module M Decoder , the M Decoder Contains 1 upsampling layer, 1 transposed convolution layer, 1 ReLU activation function layer, 1 batch normalization layer, and 1 Softmax activation function layer;
[0071] Step 4. Input the image training set T with enhanced morphological features morph And edge feature enhanced image training set T edge , for the multi-modal multi-level feature fusion learning network N seg Conduct training;
[0072] Step 4.1 T morph Each image I 1 Input morphological feature extraction subnetwork N morph to process;
[0073] Step 4.1.1 Image I 1 Divide into (W×H) / (4×4) image blocks, each image block is 4×4 in size, and get an image block sequence
[0074] Step 4.1.2 Using the Transformer encoder module right Processing is performed to obtain the morphological feature map at the first scale
[0075] Step 4.1.3 Using the Transformer encoder module right Processing is performed to obtain the morphological feature map at the second scale
[0076] Step 4.1.4 Using the Transformer encoder module right Processing is performed to obtain the morphological feature map at the third scale
[0077] Step 4.2 Tedge Each image I 2 Input edge feature extraction subnetwork N edge to process;
[0078] Step 4.2.1 Image I 2 Divide into (W×H) / (4×4) image blocks, each image block is 4×4 in size, and get an image block sequence
[0079] Step 4.2.2 Using the Transformer encoder module right Processing is performed to obtain the edge feature map at the first scale
[0080] Step 4.2.3 Using the Transformer encoder module right Processing is performed to obtain the edge feature map at the second scale
[0081] Step 4.2.4 Using the Transformer encoder module right Processing is performed to obtain the edge feature map at the third scale
[0082] Step 4.3: Feature maps at different scales and Input feature fusion subnetwork N fusion to process;
[0083] Step 4.3.1 Utilize cross-modal feature enhancement module M cfe , according to formula (3) to formula (5), calculate the feature weight λ at different scales 1 , 2 and λ 3 ;
[0084]
[0085] The lambda 1 represents the feature weight at the first scale, λ 2 represents the feature weight at the second scale, λ 3 represents the feature weight at the third scale, σ(·) represents the Sigmoid activation function, Conv T (·) represents a series of operations consisting of a convolution layer with a convolution kernel size of 3×3, a batch normalization layer, and a ReLU activation function layer. Concat(·) represents a concatenation operation.
[0086] Step 4.3.2 According to formula (6) and formula (7), the feature map and Perform weighted operation to obtain the weighted feature map and
[0087]
[0088] i∈{1,2,3},λ i represents the feature weight at the i-th scale, represents the morphological feature map at the i-th scale, represents the edge feature map at the i-th scale, represents the weighted morphological feature map at the i-th scale, Represents the weighted edge feature map at the i-th scale;
[0089] Step 4.3.3 Using the Transformer encoder module According to formula (8), the morphological feature maps and edge feature maps at different scales are processed to obtain the fused feature map and
[0090]
[0091] Said represents the fusion feature map at the i-th scale, They represent the fused feature map at the first scale, the fused feature map at the second scale, and the fused feature map at the third scale respectively;
[0092] Step 4.3.4 Enhance the fusion module M using cross-modal features cffu According to formula (9), the unweighted edge feature map, morphological feature map and fusion feature map are connected to obtain the feature map Then, according to formula (10), the fusion feature map is calculated
[0093]
[0094] Said represents the fused connection feature map at the i-th scale, represents the fusion feature map at the i-th scale, represents the matrix Hadamard product operation;
[0095] Step 4.4 Use the segmentation subnetwork N segHead Fusion feature map Processing is performed to obtain the final segmentation result;
[0096] Step 4.4.1: Calculate the upsampled mapping feature map according to formula (11) and formula (12);
[0097]
[0098] And order Said represents the upsampled mapping feature map at the first scale, represents the upsampled mapping feature map at the second scale, Represents the upsampled mapping feature map at the third scale;
[0099] Step 4.4.2: According to formula (13), the upsampled mapping feature map is connected with the final fusion feature map;
[0100]
[0101] Said represents the composite connection feature map at the i-th scale, Represents the upsampled mapping feature map at the i-th scale;
[0102] Step 4.4.3 Composite connection feature graph The convolution operation with a kernel size of 3×3, the ReLU activation operation, and the batch normalization operation are performed in sequence to obtain the output feature map O i , the O i Represents the output feature map at the i-th scale;
[0103] Step 4.4.4: transform the feature map O i Input the Softmax activation function to obtain the category probability of each pixel and the lesion segmentation result Figure I segment , and the trained multi-modal multi-level feature fusion learning network N seg ;
[0104] Step 5. Input the breast tumor MRI image J to be processed and use the trained deep convolutional neural network N seg Process J and output the lesion segmentation result image J output .
[0105] In order to verify the effectiveness of the present invention, Table 1 shows the quantitative evaluation results of breast tumor segmentation of the present invention and the BTD-ELM method, ODADA method, SharpU-Net method, and UNeXt method on clinical data sets, and the objective evaluation indicators include Dice, Jaccard, Precision, and 95HD. From the results in Table 1, it can be seen that the breast tumor segmentation performance of the present invention has been significantly improved compared with several other typical methods.
[0106] Table 1 Comparison of segmentation results of the present invention with other typical methods
[0107]
[0108] Figure 4 The subjective comparison of the segmentation results of 8 breast tumor images by the present invention and the BTD-ELM method, ODADA method, Sharp U-Net method, and UNeXt method is further given. Figure 4 It can be seen that the BTD-ELM method cannot resist the interference of fuzzy and irregular boundaries of the tumor lesion area, and mis-segmentation is more common; the ODADA method cannot accurately capture the edge contour of the tumor lesion area; the Sharp U-Net method can locate the tumor lesion area more accurately in most cases, but it is not robust enough and may cause mis-segmentation in specific cases; the UNeXt method has the phenomenon of mis-segmentation of healthy tissue areas; in contrast, the present invention effectively suppresses the interference of fuzzy and irregular boundaries of the tumor lesion area and can accurately locate the tumor lesion area.
Claims
1. A breast tumor segmentation method based on a multi-modal multi-level feature fusion learning network, characterized in that Follow these steps: Step 1. Input any number of four single-modality breast tumor MRI images to form a training set T, wherein the four single-modality MRI images include DCE images, DWI images, T1-weighted images, and T2-weighted images; Step 2: Preprocess each breast tumor MRI image I in the training set to obtain an image training set T with enhanced morphological features. morph And edge feature enhanced image training set T edge , let the height of image I be H and the width be W; Step 2.1: Use the convolution kernel H shown in formula (1) to perform convolution operation on image I to enhance the contrast of the edge of the breast tumor lesion in image I, thereby obtaining image I with enhanced morphological features. 1 , forming the image training set T with enhanced morphological features morph ; Step 2.2 Let the value of the pixel at coordinate (x, y) in image I be f(x, y), and use the gradient difference method to calculate the horizontal gradient value Grad of f(x, y) respectively. x (x, y), vertical gradient value Grad y (x, y), main diagonal gradient value Grad d1 (x, y), sub-diagonal gradient value Grad d2 (x, y), and then calculate the weighted gradient value Grad(x, y) of f(x, y) according to formula (2) to obtain the image I after edge feature enhancement 2 , forming the edge feature enhanced image training set T edge ; Grad(x,y)=ω1×Grad x (x,y)+ω2×Grad y (x,y)+ω3×Grad d1 (x,y)+ω4×Grad d2 (x,y)(2) The 1≤x≤W, 1≤y≤H, ω1, ω2, ω3 and ω4 are preset weight constants; Step 3. Establish and initialize the multimodal multi-level feature fusion learning network N seg , including 1 morphological feature extraction sub-network N morph , 1 edge feature extraction sub-network N edge , 1 feature fusion sub-network N fusion , 1 segmentation sub-network N segHead ; Step 3.1 Establish and initialize the morphological feature extraction subnetwork N morph , including 3 Transformer encoder modules, namely Said The input data dimension is W×H, and the output data dimension is W / 4×H / 4×C1. The input data dimension is W / 4×H / 4×C1, and the output data dimension is W / 8×H / 8×C2. The input data dimension is W / 8×H / 8×C2, and the output data dimension is W / 16×H / 16×C3, where C1, C2, and C3 are preset constants; Step 3.2 Establish and initialize the edge feature extraction subnetwork N edge , including 3 Transformer encoder modules, namely Said The input data dimension is W×H, and the output data dimension is W / 4×H / 4×C1. The input data dimension is W / 4×H / 4×C1, and the output data dimension is W / 8×H / 8×C2. The input data dimension is W / 8×H / 8×C2, and the output data dimension is W / 16×H / 16×C3; Step 3.3 Establish and initialize the feature fusion subnetwork N fusion , including a cross-modal feature enhancement module M cfe , 1 Transformer encoder module 1 cross-modal feature enhancement fusion module M cffu , the cross-modal feature enhancement module M cfe Contains 1 Sigmoid activation function layer, 1 convolution layer with a convolution kernel size of 3×3, 1 batch normalization layer, and 1 ReLU activation function layer; Step 3.4 Establish and initialize the segmentation subnetwork N segHead , including a sequence decoder module M Decoder , the M Decoder Contains 1 upsampling layer, 1 transposed convolution layer, 1 ReLU activation function layer, 1 batch normalization layer, and 1 Softmax activation function layer; Step 4. Input the image training set T with enhanced morphological features morph And edge feature enhanced image training set T edge , for the multi-modal multi-level feature fusion learning network N seg Conduct training; Step 4.1 T morph Each image I 1 Input morphological feature extraction subnetwork N morph to process; Step 4.1.1 Image I 1 Divide into (W×H) / (4×4) image blocks, each image block is 4×4 in size, and get an image block sequence Step 4.1.2 Using the Transformer encoder module right Processing is performed to obtain the morphological feature map at the first scale Step 4.1.3 Using the Transformer encoder module right Processing is performed to obtain the morphological feature map at the second scale Step 4.1.4 Using the Transformer encoder module right Processing is performed to obtain the morphological feature map at the third scale Step 4.2 T edge Each image I 2 Input edge feature extraction subnetwork N edge to process; Step 4.2.1 Image I 2 Divide into (W×H) / (4×4) image blocks, each image block is 4×4 in size, and get an image block sequence Step 4.2.2 Using the Transformer encoder module right Processing is performed to obtain the edge feature map at the first scale Step 4.2.3 Using the Transformer encoder module right Processing is performed to obtain the edge feature map at the second scale Step 4.2.4 Using the Transformer encoder module right Processing is performed to obtain the edge feature map at the third scale Step 4.3: Feature maps at different scales and Input feature fusion subnetwork N fusion to process; Step 4.3.1 Utilize cross-modal feature enhancement module M cfe , according to formula (3) to formula (5), calculate the feature weights λ1, λ2 and λ3 at different scales; The λ1 represents the feature weight at the first scale, λ2 represents the feature weight at the second scale, λ3 represents the feature weight at the third scale, σ(·) represents the Sigmoid activation function, Conv T (·) represents a series of operations consisting of a convolution layer with a convolution kernel size of 3×3, a batch normalization layer, and a ReLU activation function layer. Concat(·) represents a concatenation operation. Step 4.3.2 According to formula (6) and formula (7), the feature map and Perform weighted operation to obtain the weighted feature map and i∈{1,2,3},λ i represents the feature weight at the i-th scale, represents the morphological feature map at the i-th scale, represents the edge feature map at the i-th scale, represents the weighted morphological feature map at the i-th scale, Represents the weighted edge feature map at the i-th scale; Step 4.3.3 Using the Transformer encoder module According to formula (8), the morphological feature maps and edge feature maps at different scales are processed to obtain the fused feature map and Said represents the fusion feature map at the i-th scale, They represent the fused feature map at the first scale, the fused feature map at the second scale, and the fused feature map at the third scale respectively; Step 4.3.4 Enhance the fusion module M using cross-modal features cffu According to formula (9), the unweighted edge feature map, morphological feature map and fusion feature map are connected to obtain the feature map Then according to formula (10), we get the fusion feature map Said represents the fused connection feature map at the i-th scale, represents the fusion feature map at the i-th scale, represents the matrix Hadamard product operation; Step 4.4 Use the segmentation subnetwork N segHead Fusion feature map Processing is performed to obtain the final segmentation result; Step 4.4.1: Calculate the upsampled mapping feature map according to formula (11) and formula (12); And order Said represents the upsampled mapping feature map at the first scale, represents the upsampled mapping feature map at the second scale, Represents the upsampled mapping feature map at the third scale; Step 4.4.2: According to formula (13), the upsampled mapping feature map is connected with the final fusion feature map; Said represents the composite connection feature map at the i-th scale, Represents the upsampled mapping feature map at the i-th scale; Step 4.4.3 Composite connection feature graph The convolution operation with a kernel size of 3×3, the ReLU activation operation, and the batch normalization operation are performed in sequence to obtain the output feature map O i , the O i Represents the output feature map at the i-th scale; Step 4.4.4: transform the feature map O i Input the Softmax activation function to obtain the category probability of each pixel and the lesion segmentation result Figure I segment And the trained multi-modal multi-level feature fusion learning network N seg ; Step 5. Input the breast tumor MRI image J to be processed and use the trained deep convolutional neural network N seg Process J and output the lesion segmentation result image J output .