Endometrial cancer diagnosis method based on multi-mode flexible classification network
By adopting a multimodal flexible classification network in the diagnosis of endometrial cancer, the parallel structure of student subnet, teacher subnet and cross-modal feature fusion enhancer network is solved, and high-precision endometrial cancer diagnosis is achieved.
Patent Information
- Application Number
- CN202510031318.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art has problems such as difficulty in character fusion, high requirements for multimodal information alignment, and adaptive selection of feature fusion stage in the diagnosis of endometrial cancer, and has failed to effectively discover multimodal complementary features and cross-modal feature alignment.
The endometrial cancer diagnosis method based on multimodal flexible classification network is adopted, and the multi-scale feature extraction and adaptive feature weight adjustment of multimodal data are realized through the parallel structure of student subnet, teacher subnet and cross-modal feature fusion enhancer network.
It improves the recognition accuracy of deep neural networks for cancerous areas, enhances feature extraction capabilities and classification performance, reduces the alignment accuracy requirements, and improves the accuracy and robustness of diagnosis.
Smart Images

Figure CN120108689A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical image processing, and in particular to a method for diagnosing endometrial cancer based on a multimodal flexible classification network. Background Art
[0002] In recent years, endometrial cancer, as a common gynecological malignancy, has shown an increasing incidence and mortality rate year by year, posing a serious threat to women's safety and quality of life. Endometrial cancer originates from the endometrium in the early stages, but cancer cells spread to the myometrium and surrounding tissues as the disease progresses. Therefore, early diagnosis and treatment of endometrial cancer are particularly important. However, the clinical manifestations of endometrial cancer often lack specificity and are easily confused with other gynecological diseases, which undoubtedly increases the difficulty and risk of its diagnosis.
[0003] Magnetic resonance imaging (MRI) can provide doctors with three-dimensional visualization data of human tissues and organs, and has a high soft tissue contrast. It has become one of the indispensable and important tools for diagnosing endometrial cancer. Doctors can accurately determine the location, size and morphological characteristics of the tumor through MRI, and can evaluate the malignancy of the tumor and predict the treatment response. Therefore, it plays an important role in the screening, diagnosis, staging and follow-up of breast tumors. However, single-modality MRI images can often only reflect one aspect of the characteristics of endometrial cancer, and there are significant limitations in the diagnosis of endometrial cancer. For example, T1 images mainly enhance the signals of blood vessels and muscle organs, while DWI sequence images focus on reflecting the signal intensity changes in the lesion area. Given the complex and diverse pathological changes of endometrial cancer, it may produce characteristics involving multiple aspects, which cannot be fully captured and accurately interpreted in single-modality images, which may lead to missed diagnosis or misdiagnosis of the disease. In this case, computer-assisted analysis of endometrial tissue using multimodal MRI images and deep neural network technology has important clinical value for improving the diagnosis and treatment efficiency of endometrial cancer.
[0004] Although multimodal MRI combined with deep neural network technology has shown great potential in the diagnosis of endometrial cancer, there are still three key technical challenges that have not been effectively resolved and still require further technical research and development and clinical verification.
[0005] First, feature fusion is difficult. The imaging objectives and significance of each modality are different, which makes different modalities of MRI (such as T2 weighted, DWI dynamic contrast enhancement) have significant feature heterogeneity in imaging features. It is necessary to design a suitable feature fusion strategy to effectively fuse the information of these different modalities.
[0006] Second, the alignment of multimodal information has high requirements. Especially when the patient's position changes, aligning image data from different modalities is a complex task. Only by ensuring the spatial consistency and alignment accuracy of multimodal image data can the fused feature information be guaranteed to have clinical value.
[0007] Third, the feature fusion stage requires adaptive selection. Early fusion methods directly combine images of different modalities at the input stage, which may lose some modality details; late fusion trains the model for each modality separately and then fuses the results, which may lead to inconsistent model output results.
[0008] In summary, there is currently no diagnostic method for endometrial cancer that combines multimodal MRI images with deep neural networks, which can effectively explore multimodal complementary features and has the ability to align cross-modal features and adaptively select in the feature fusion stage. Summary of the invention
[0009] The present invention aims to solve the above-mentioned technical problems existing in the prior art and provides a method for diagnosing endometrial cancer based on a multimodal flexible classification network.
[0010] The technical solution of the present invention is: a method for diagnosing endometrial cancer based on a multimodal flexible classification network, which is carried out in the following steps:
[0011] Step 1. Input any number of four single-modality endometrial MRI images to form a training set T, wherein the four single-modality MRI images include ADC images, DWI images, T1-weighted images, and T2-weighted images;
[0012] Step 2. Combine the T1-weighted images in the training set T into the superior modality training set T SU , and the ADC images, DWI images, and T2-weighted images in T form the inferior modality training set T IN ;
[0013] Step 3. Establish and initialize the multimodal flexible classification network, including 1 student sub-network N student , 1 teacher sub-network N teacher , 1 cross-modal feature fusion enhancement sub-network N fusion , 1 classification prediction subnetwork N pred ;
[0014] Step 3.1 Establish and initialize the student subnetwork N student , including 1 pre-trained AlexNet module 1 convolution module 1 convolution module 1 fully connected module 1 fully connected module
[0015] The convolution module It contains 1 convolution layer consisting of 512 convolution kernels of size 3×3 and edge padding mode 1, 1 ReLU activation function layer, and 1 global pooling layer with a window size of 2×2;
[0016] The convolution module It contains 1 convolutional layer consisting of 512 convolutional kernels of size 3×3 and edge padding mode 1, 1 ReLU activation function layer, and 1 global maximum pooling layer with window size 2×2 and stride 2;
[0017] The fully connected module Contains 1 fully connected layer with 4096 neurons and 1 ReLU activation function layer;
[0018] The fully connected module Contains 1 fully connected layer with 64 neurons;
[0019] Step 3.2 Establish and initialize the teacher sub-network N teacher , including 1 pre-trained AlexNet module 1 convolution module 1 convolution module 1 fully connected module 1 fully connected module
[0020] The convolution module It contains 1 convolution layer consisting of 512 convolution kernels of size 3×3 and edge padding mode 1, 1 ReLU activation function layer, and 1 global pooling layer with a window size of 2×2;
[0021] The convolution module It contains 1 convolution layer consisting of 512 convolution kernels of size 3×3 and edge padding mode 1, 1 ReLU activation function layer, and 1 global maximum pooling layer with window size 2×2 and stride 2;
[0022] The fully connected module Contains 1 fully connected layer with 4096 neurons and 1 ReLU activation function layer;
[0023] The fully connected module Contains 1 fully connected layer with 64 neurons;
[0024] Step 3.3 Establish and initialize the cross-modal feature fusion enhancement sub-network N fusion , including 3 convolution modules, namely
[0025] Said It contains 1 convolution layer consisting of 64 convolution kernels of size 3×3 and edge padding mode 1, and 1 Sigmoid activation function layer, where the stride of each convolution kernel is 1 and the edge padding mode is 1;
[0026] Said Contains 1 convolutional layer consisting of 64 convolution kernels of size 3×3 and edge padding mode 1. Each convolution kernel has a stride of 1 and an edge padding mode of 1.
[0027] Said Contains 1 convolutional layer consisting of 64 convolution kernels of size 3×3 and edge padding mode 1. Each convolution kernel has a stride of 1 and an edge padding mode of 1.
[0028] Step 3.4 Establish and initialize the classification prediction subnetwork N pred , including 1 loss function module And 3 layers of Softmax activation function layers, respectively
[0029] Step 4. Input the superior modality training set T SU and the inferior modality training set T IN , train the multimodal flexible classification network;
[0030] Step 4.1 T IN Each image I IN Input student subnetwork N student to process;
[0031] Step 4.1.1 Using the pre-trained AlexNet module to I IN Processing to obtain feature map
[0032] Step 4.1.2 Using the convolution module right Processing to obtain feature map
[0033] Step 4.1.3 Using the convolutional module right Processing to obtain feature map
[0034] Step 4.1.4 Using fully connected modules right Processing to obtain feature map
[0035] Step 4.1.5 Using fully connected modules right Processing to obtain feature map
[0036] Step 4.2 T SU Each image I SU Input teacher subnetwork N teacher to process;
[0037] Step 4.2.1 Using the pre-trained AlexNet module to I SU Processing to obtain feature map
[0038] Step 4.2.2 Using the convolution module right Processing to obtain feature map
[0039] Step 4.2.3 Using the convolution module right Processing to obtain feature map
[0040] Step 4.2.4 Using fully connected modules right Processing to obtain feature map
[0041] Step 4.2.5 Using fully connected modules right Processing to obtain feature map
[0042] Step 4.3 Enhance the sub-network N using cross-modal feature fusion fusion Fuse the feature maps calculated by the student sub-network and the teacher sub-network;
[0043] Step 4.3.1 Feature map and feature map Perform the connection operation to obtain the feature map
[0044] Step 4.3.2 Using the convolution module right Processing to obtain feature map
[0045] Step 4.3.3: Feature map With feature map Perform Hadamard product to get the feature map
[0046] Step 4.3.4: Feature map With feature map Perform Hadamard product to get the feature map
[0047] Step 4.3.5: Feature map With feature map Add together to get the feature map
[0048] Step 4.3.6 Using the convolution module right Processing to obtain feature map
[0049] Step 4.3.7: Feature map Perform the connection operation to obtain the feature map
[0050] Step 4.3.8 Using the convolution module right Processing to obtain feature map
[0051] Step 4.3.9: Feature map Perform the connection operation to obtain the feature map
[0052] Step 4.4 Use the classification prediction subnetwork N pred Calculate the category probability of each pixel to obtain a trained multimodal flexible classification network;
[0053] Step 4.4.1 Use Softmax activation function layer For feature maps Calculate and get the predicted label distribution p of the student sub-network student ;
[0054] Step 4.4.2 Use Softmax activation function layer For feature maps Calculate and get the predicted label distribution p of the teacher sub-network teacher ;
[0055] Step 4.4.3 Using Softmax activation function layer For feature maps Calculate and obtain the predicted label distribution p of the cross-modal feature fusion enhanced sub-network fusion ;
[0056] Step 4.4.4 Calculate p according to formula (1)-formula (3) student With the true label distribution p ideal The mixed loss function value Lstudent ;
[0057]
[0058] The n represents the image I IN The number of pixels, p student (i) indicates p student The predicted value of the i-th pixel in ideal (i) represents the true label distribution p ideal The predicted value of the i-th pixel in , represents the Dice loss function value of the student sub-network, Represents the cross entropy loss function value of the student sub-network;
[0059] Step 4.4.5 Calculate p according to formula (4)-formula (6) teacher With the true label distribution p ideal The mixed loss function value L teacher ;
[0060]
[0061] The p teacher (i) indicates p teacher The predicted value of the i-th pixel in , represents the Dice loss function value of the teacher sub-network, Represents the cross entropy loss function value of the teacher sub-network;
[0062] Step 4.4.6 Calculate p according to formula (7)-formula (9) fusion With the true label distribution p ideal The mixed loss function value L fusion ;
[0063]
[0064]
[0065] The p fusion (i) indicates p fusion The predicted value of the i-th pixel in , represents the Dice loss function value of the cross-modal feature fusion enhanced sub-network, Represents the cross entropy loss function value of the cross-modal feature fusion enhancement sub-network;
[0066] Step 4.4.7 Calculate the loss function value L of the training set T according to formula (10): total , and then use the back-propagation algorithm to iteratively update the network parameters to obtain a trained multi-modal flexible classification network;
[0067] Ltotal =L student +L teacher +L fusion (10)
[0068] Step 5. Input the endometrial MRI image J to be processed, use the trained multimodal flexible classification network to process J, and output the lesion classification result J output .
[0069] Compared with the prior art, the present invention has the following advantages: First, the present invention proposes a cross-modal feature enhancement module, which uses a hierarchical fusion strategy to fully explore the heterogeneous features and complementary information between different modalities, can enhance the model's feature extraction ability and classification performance, obtain key pathological features, and then improve the recognition accuracy of deep neural networks for cancerous areas, solving the problem of difficult feature fusion that is common in prior art. Second, the knowledge distillation strategy is introduced, and the teacher sub-network is used to learn deep features from superior modality images, and guide the student sub-network to extract complementary and useful features for endometrial cancer diagnosis from inferior modality images, which not only optimizes the network model's ability to extract features from inferior modality images, ensuring the accuracy and robustness of the prediction results, but also reduces the network model's requirements for the alignment accuracy of multimodal data, improving the robustness and clinical usability of the network model. Third, the present invention introduces a multimodal flexible classification strategy, using three parallel network structures, namely, a student subnetwork, a teacher subnetwork, and a cross-modal feature fusion enhancement subnetwork, to extract multi-scale features of multimodal data respectively. On this basis, the feature weights of different modal data and their fusion strengths are adaptively adjusted, which can more effectively extract the deep-level features of multimodal endometrial MRI images, thereby maximizing the diagnostic advantages of multimodal MRI images for endometrial cancer. In summary, the present invention has the characteristics of strong feature perception ability, excellent multimodal feature mining performance, and high multimodal feature fusion efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 It is a schematic diagram of an MRI image and a lesion location according to an embodiment of the present invention.
[0071] Figure 2 4 is a diagram of a neural network structure according to an embodiment of the present invention.
[0072] Figure 3 4 is a graph showing the experimental results of LVSI infiltrative property prediction according to an embodiment of the present invention.
[0073] Figure 4 4 is a graph showing the experimental results of pathological grading and LVSI invasiveness prediction according to an embodiment of the present invention. DETAILED DESCRIPTION
[0074] The present invention provides a method for classifying breast tumors using a multi-modal multi-level feature fusion learning network, which is characterized by being performed in the following steps:
[0075] Step 1. Enter any number of Figure 1 The four single-modality endometrial MRI images shown constitute a training set T, wherein the four single-modality MRI images include a. T1-weighted image, b. T2-weighted image, c. ADC image, and d. DWI image;
[0076] Step 2. Combine the T1-weighted images in the training set T into the superior modality training set T SU , and the ADC images, DWI images, and T2-weighted images in T form the inferior modality training set T IN ;
[0077] Step 3. Create and initialize Figure 2 The multimodal flexible classification network shown includes 1 student subnetwork N student , 1 teacher sub-network N teacher , 1 cross-modal feature fusion enhancement sub-network N fusion , 1 classification prediction subnetwork N pred ;
[0078] Step 3.1 Establish and initialize the student subnetwork N student , including 1 pre-trained AlexNet module 1 convolution module 1 convolution module 1 fully connected module 1 fully connected module
[0079] The convolution module It contains 1 convolution layer consisting of 512 convolution kernels of size 3×3 and edge padding mode 1, 1 ReLU activation function layer, and 1 global pooling layer with a window size of 2×2;
[0080] The convolution module It contains 1 convolutional layer consisting of 512 convolutional kernels of size 3×3 and edge padding mode 1, 1 ReLU activation function layer, and 1 global maximum pooling layer with window size 2×2 and stride 2;
[0081] The fully connected module Contains 1 fully connected layer with 4096 neurons and 1 ReLU activation function layer;
[0082] The fully connected module Contains 1 fully connected layer with 64 neurons;
[0083] Step 3.2 Establish and initialize the teacher sub-network N teacher , including 1 pre-trained AlexNet module 1 convolution module 1 convolution module 1 fully connected module 1 fully connected module
[0084] The convolution module It contains 1 convolution layer consisting of 512 convolution kernels of size 3×3 and edge padding mode 1, 1 ReLU activation function layer, and 1 global pooling layer with a window size of 2×2;
[0085] The convolution module It contains 1 convolution layer consisting of 512 convolution kernels of size 3×3 and edge padding mode 1, 1 ReLU activation function layer, and 1 global maximum pooling layer with window size 2×2 and stride 2;
[0086] The fully connected module Contains 1 fully connected layer with 4096 neurons and 1 ReLU activation function layer;
[0087] The fully connected module Contains 1 fully connected layer with 64 neurons;
[0088] Step 3.3 Establish and initialize the cross-modal feature fusion enhancement sub-network N fusion , including 3 convolution modules, namely
[0089] Said It contains 1 convolution layer consisting of 64 convolution kernels of size 3×3 and edge padding mode 1, and 1 Sigmoid activation function layer, where the stride of each convolution kernel is 1 and the edge padding mode is 1;
[0090] Said Contains 1 convolutional layer consisting of 64 convolution kernels of size 3×3 and edge padding mode 1. Each convolution kernel has a stride of 1 and an edge padding mode of 1.
[0091] Said Contains 1 convolutional layer consisting of 64 convolution kernels of size 3×3 and edge padding mode 1. Each convolution kernel has a stride of 1 and an edge padding mode of 1.
[0092] Step 3.4 Establish and initialize the classification prediction subnetwork N pred , including 1 loss function module And 3 layers of Softmax activation function layers, respectively
[0093] Step 4. Input the superior modality training set T SU and the inferior modality training set T IN , train the multimodal flexible classification network;
[0094] Step 4.1 T IN Each image I IN Input student subnetwork N student to process;
[0095] Step 4.1.1 Using the pre-trained AlexNet module to I IN Processing to obtain feature map
[0096] Step 4.1.2 Using the convolution module right Processing to obtain feature map
[0097] Step 4.1.3 Using the convolutional module right Processing to obtain feature map
[0098] Step 4.1.4 Using fully connected modules right Processing to obtain feature map
[0099] Step 4.1.5 Using fully connected modules right Processing to obtain feature map
[0100] Step 4.2 T SU Each image I SU Input teacher subnetwork N teacher to process;
[0101] Step 4.2.1 Using the pre-trained AlexNet module to I SU Processing to obtain feature map
[0102] Step 4.2.2 Using the convolution module right Processing to obtain feature map
[0103] Step 4.2.3 Using the convolution module right Processing to obtain feature map
[0104] Step 4.2.4 Using the fully connected module right Processing to obtain feature map
[0105] Step 4.2.5 Using fully connected modules right Processing to obtain feature map
[0106] Step 4.3 Enhance the sub-network N using cross-modal feature fusion fusion Fuse the feature maps calculated by the student sub-network and the teacher sub-network;
[0107] Step 4.3.1 Feature map and feature map Perform the connection operation to obtain the feature map
[0108] Step 4.3.2 Using the convolution module right Processing to obtain feature map
[0109] Step 4.3.3: Feature map With feature map Perform Hadamard product to get the feature map
[0110] Step 4.3.4: Feature map With feature map Perform Hadamard product to get the feature map
[0111] Step 4.3.5: Feature map With feature map Add together to get the feature map
[0112] Step 4.3.6 Using the convolution module right Processing to obtain feature map
[0113] Step 4.3.7: Feature map Perform the connection operation to obtain the feature map
[0114] Step 4.3.8 Using the convolution module right Processing to obtain feature map
[0115] Step 4.3.9: Feature map Perform the connection operation to obtain the feature map
[0116] Step 4.4 Use the classification prediction subnetwork N pred Calculate the category probability of each pixel to obtain a trained multimodal flexible classification network. In this embodiment, the stochastic gradient descent method is selected as the optimizer, and the initial learning rate is set to 0.0001, the maximum number of iterations is set to 20, and the batch size is set to 128;
[0117] Step 4.4.1 Use Softmax activation function layer For feature maps Calculate and get the predicted label distribution p of the student sub-network student ;
[0118] Step 4.4.2 Use Softmax activation function layer For feature maps Calculate and get the predicted label distribution p of the teacher sub-network teacher ;
[0119] Step 4.4.3 Using Softmax activation function layer For feature maps Calculate and obtain the predicted label distribution p of the cross-modal feature fusion enhanced sub-network fusion ;
[0120] Step 4.4.4 Calculate p according to formula (1)-formula (3) student With the true label distribution p ideal The mixed loss function value L student ;
[0121]
[0122] The n represents the image I IN The number of pixels, p student (i) indicates p student The predicted value of the i-th pixel in ideal (i) represents the true label distribution p ideal The predicted value of the i-th pixel in , represents the Dice loss function value of the student sub-network, Represents the cross entropy loss function value of the student sub-network;
[0123] Step 4.4.5 Calculate p according to formula (4)-formula (6) teacher With the true label distribution p idealThe mixed loss function value L teacher ;
[0124]
[0125] The p teacher (i) indicates p teacher The predicted value of the i-th pixel in , represents the Dice loss function value of the teacher sub-network, Represents the cross entropy loss function value of the teacher sub-network;
[0126] Step 4.4.6 Calculate p according to formula (7)-formula (9) fusion With the true label distribution p ideal The mixed loss function value L fusion ;
[0127]
[0128] The p fusion (i) indicates p fusion The predicted value of the i-th pixel in , represents the Dice loss function value of the cross-modal feature fusion enhanced sub-network, Represents the cross entropy loss function value of the cross-modal feature fusion enhancement sub-network;
[0129] Step 4.4.7 Calculate the loss function value L of the training set T according to formula (10): total , and then use the back-propagation algorithm to iteratively update the network parameters to obtain a trained multi-modal flexible classification network;
[0130] L total =L student +L teacher +L fusion (10)
[0131] Step 5. Input the endometrial MRI image J to be processed, use the trained multimodal flexible classification network to process J, and output the lesion classification result J output .
[0132] In order to verify the effectiveness of the present invention, the present invention collected clinical MRI image data of 297 endometrial cancer patients from the Second Affiliated Hospital of Dalian Medical University to form a data set, including ADC images, DWI images, T1-weighted images, and T2-weighted images. According to the patient's cancer grade and LVSI positivity, the data set is subdivided into 5 categories, including 97 cases of first-level negative, 32 cases of second-level positive, 116 cases of second-level negative, 30 cases of third-level positive, and 22 cases of third-level negative. The data set and lesion area samples are shown in the following figure. Figure 1 As shown, Figure 1 (a) is the image and lesion area under T1 mode. Figure 1 (b) is the image and lesion area under ADC mode. Figure 1 (c) is the image and lesion area under T2 mode. Figure 1 (d) is the image and its lesion area under the DWI modality. Further, the data set is divided into a training set and a test set, where the training set accounts for 70% and the test set accounts for 30%.
[0133] Figure 3 Shown are the experimental results of LVSI infiltrative prediction of the present invention; Figure 4 The figure shows the experimental results of the pathological grading and LVSI infiltration prediction of the present invention, wherein + represents positive LVSI infiltration, - represents positive LVSI infiltration, and 1- represents a pathological grade of one and negative LVSI infiltration. There are no patients with a pathological grade of one and positive LVSI infiltration in the clinic, so there is no 1+ category. Figure 3 and Figure 4 It can be seen that the present invention can accurately identify the type and stage of endometrial cancer, and provide important decision-making support for clinicians.
Claims
1. A method for diagnosing endometrial cancer based on a multimodal flexible classification network, characterized in that Follow these steps: Step 1. Input any number of four single-modality endometrial MRI images to form a training set T, wherein the four single-modality MRI images include ADC images, DWI images, T1-weighted images, and T2-weighted images; Step 2. Combine the T1-weighted images in the training set T into the superior modality training set T SU , and the ADC images, DWI images, and T2-weighted images in T form the inferior modality training set T IN ; Step 3. Establish and initialize the multimodal flexible classification network, including 1 student sub-network N student , 1 teacher sub-network N teacher , 1 cross-modal feature fusion enhancement sub-network N fusion , 1 classification prediction subnetwork N pred ; Step 3.1 Establish and initialize the student subnetwork N student , including 1 pre-trained AlexNet module 1 convolution module 1 convolution module 1 fully connected module 1 fully connected module The convolution module It contains 1 convolution layer consisting of 512 convolution kernels of size 3×3 and edge padding mode 1, 1 ReLU activation function layer, and 1 global pooling layer with a window size of 2×2; The convolution module It contains 1 convolutional layer consisting of 512 convolutional kernels of size 3×3 and edge padding mode 1, 1 ReLU activation function layer, and 1 global maximum pooling layer with window size 2×2 and stride 2; The fully connected module Contains 1 fully connected layer with 4096 neurons and 1 ReLU activation function layer; The fully connected module Contains 1 fully connected layer with 64 neurons; Step 3.2 Establish and initialize the teacher sub-network N teacher , including 1 pre-trained AlexNet module 1 convolution module 1 convolution module 1 fully connected module 1 fully connected module The convolution module It contains 1 convolution layer consisting of 512 convolution kernels of size 3×3 and edge padding mode 1, 1 ReLU activation function layer, and 1 global pooling layer with a window size of 2×2; The convolution module It contains 1 convolution layer consisting of 512 convolution kernels of size 3×3 and edge padding mode 1, 1 ReLU activation function layer, and 1 global maximum pooling layer with window size 2×2 and stride 2; The fully connected module Contains 1 fully connected layer with 4096 neurons and 1 ReLU activation function layer; The fully connected module Contains 1 fully connected layer with 64 neurons; Step 3.3 Establish and initialize the cross-modal feature fusion enhancement sub-network N fusion , including 3 convolution modules, namely Said It contains 1 convolution layer consisting of 64 convolution kernels of size 3×3 and edge padding mode 1, and 1 Sigmoid activation function layer, where the stride of each convolution kernel is 1 and the edge padding mode is 1; Said Contains 1 convolutional layer consisting of 64 convolution kernels of size 3×3 and edge padding mode 1. Each convolution kernel has a stride of 1 and an edge padding mode of 1. Said Contains 1 convolutional layer consisting of 64 convolution kernels of size 3×3 and edge padding mode 1. Each convolution kernel has a stride of 1 and an edge padding mode of 1. Step 3.4 Establish and initialize the classification prediction subnetwork N pred , including 1 loss function module And 3 layers of Softmax activation function layers, respectively Step 4. Input the superior modality training set T SU and the inferior modality training set T IN , train the multimodal flexible classification network; Step 4.1 T IN Each image I IN Input student subnetwork N student to process; Step 4.1.1 Using the pre-trained AlexNet module to I IN Processing to obtain feature map Step 4.1.2 Using the convolution module right Processing to obtain feature map Step 4.1.3 Using the convolutional module right Processing to obtain feature map Step 4.1.4 Using fully connected modules right Processing to obtain feature map Step 4.1.5 Using fully connected modules right Processing to obtain feature map Step 4.2 T SU Each image I SU Input teacher subnetwork N teacher to process; Step 4.2.1 Using the pre-trained AlexNet module to I SU Processing to obtain feature map Step 4.2.2 Using the convolution module right Processing to obtain feature map Step 4.2.3 Using the convolution module right Processing to obtain feature map Step 4.2.4 Using the fully connected module right Processing to obtain feature map Step 4.2.5 Using fully connected modules right Processing to obtain feature map Step 4.3 Enhance the sub-network N using cross-modal feature fusion fusion Fuse the feature maps calculated by the student sub-network and the teacher sub-network; Step 4.3.1 Feature map and feature map Perform the connection operation to obtain the feature map Step 4.3.2 Using the convolution module right Processing to obtain feature map Step 4.3.3: Feature map With feature map Perform Hadamard product to get the feature map Step 4.3.4: Feature map With feature map Perform Hadamard product to get the feature map Step 4.3.5: Feature map With feature map Add together to get the feature map Step 4.3.6 Using the convolution module right Processing to obtain feature map Step 4.3.7: Feature map Perform the connection operation to obtain the feature map Step 4.3.8 Using the convolution module right Processing to obtain feature map Step 4.3.9: Feature map Perform the connection operation to obtain the feature map Step 4.4 Use the classification prediction subnetwork N pred Calculate the category probability of each pixel to obtain a trained multimodal flexible classification network; Step 4.4.1 Use Softmax activation function layer For feature maps Calculate and get the predicted label distribution p of the student sub-network student ; Step 4.4.2 Using Softmax activation function layer For feature maps Calculate and get the predicted label distribution p of the teacher sub-network teacher ; Step 4.4.3 Using Softmax activation function layer For feature maps Calculate and obtain the predicted label distribution p of the cross-modal feature fusion enhanced sub-network fusion ; Step 4.4.4 Calculate p according to formula (1)-formula (3) student With the true label distribution p ideal The mixed loss function value L student ; The n represents the image I IN The number of pixels, p student (i) indicates p student The predicted value of the i-th pixel in ideal (i) represents the true label distribution p ideal The predicted value of the i-th pixel in , represents the Dice loss function value of the student sub-network, Represents the cross entropy loss function value of the student sub-network; Step 4.4.5 Calculate p according to formula (4)-formula (6) teacher With the true label distribution p ideal The mixed loss function value L teacher ; The p teacher (i) indicates p teacher The predicted value of the i-th pixel in , represents the Dice loss function value of the teacher sub-network, Represents the cross entropy loss function value of the teacher sub-network; Step 4.4.6 Calculate p according to formula (7)-formula (9) fusion With the true label distribution p ideal The mixed loss function value L fusion ; The p fusion (i) indicates p fusion The predicted value of the i-th pixel in , represents the Dice loss function value of the cross-modal feature fusion enhanced sub-network, Represents the cross entropy loss function value of the cross-modal feature fusion enhancement sub-network; Step 4.4.7 Calculate the loss function value L of the training set T according to formula (10): total , and then use the back-propagation algorithm to iteratively update the network parameters to obtain a trained multi-modal flexible classification network; L total =L student +L teacher +L fusion (10) Step 5. Input the endometrial MRI image J to be processed, use the trained multimodal flexible classification network to process J, and output the lesion classification result J output .
Citation Information
Cited By
Multi-modal medical image analysis system and method
CN121601273A